Skip to contents

Motivation

When a new version of a dataset arrives, the first thing we usually want to know is what changed since the last one. dfdiffs answers that with three questions:

  1. What rows are here now that weren’t here before?
  2. What rows were here before that aren’t here now?
  3. What values have been changed?

Packages

We’ll start by loading dfdiffs:

The code below loads the other packages used in this vignette:

Package functions

dfdiffs has a function for each question above, and each function comes with a pair of example datasets that show how it works. We’ll walk through each pair below.

Call structure

The four comparison functions are built from a few small helpers, and the same functions power the package’s Shiny app. I generated the call trees in this section from the package source with stackcallr (pak::pak("mjfrigaard/stackcallr")), and they only show functions defined in dfdiffs.

library(stackcallr)
call_tree_dir("R", root = "create_new_data")
call_tree_dir("R", root = "create_deleted_data")
call_tree_dir("R", root = "create_changed_data")
call_tree_dir("R", root = "create_modified_data")

create_new_data() and create_deleted_data() answer the first two questions (new rows and deleted rows), and both rely on the same four helpers: anti_join_base(), select_cols(), rename_join_col(), and create_new_column(). These helpers are written in base R, so no external comparison package is involved:

█─create_new_data
├─anti_join_base
├─select_cols
├─rename_join_col
└─create_new_column
█─create_deleted_data
├─anti_join_base
├─select_cols
├─rename_join_col
└─create_new_column

create_changed_data() and create_modified_data() answer the third question (changed values), and both build on diff_values_by_key(). diff_values_by_key() compares matched rows column by column with compare_values(), which is tolerance-aware for numeric and date-time values and label-aware for factors. Any class mismatches are recorded with class_diffs(). The call tree for create_changed_data() is shown below:

█─create_changed_data
├─select_cols
├─█─diff_values_by_key
│ ├─compare_values
│ └─class_diffs
├─rename_join_col
├─create_new_column
└─class_diffs

The Shiny app (launch_app()) wires these functions together with three modules (upload, select, and compare), and each module has a UI function and a server function. app_ui() and app_server() assemble the modules, and dfdiffs_fresh_theme() supplies the app’s theme. The code below generates the full call tree for the app:

call_tree_dir("R", root = "launch_app")
█─launch_app
├─█─app_ui
│ ├─dfdiffs_fresh_theme
│ ├─mod_upload_ui
│ ├─mod_select_ui
│ └─mod_compare_ui
└─█─app_server
  ├─█─mod_upload_server
  │ ├─█─upload_data
  │ │ └─load_flat_file
  │ ├─base_react_theme
  │ └─comp_react_theme
  ├─█─mod_select_server
  │ ├─select_cols
  │ ├─base_react_theme
  │ ├─comp_react_theme
  │ ├─info_react_theme
  │ └─█─create_join_column
  │   └─select_cols
  └─█─mod_compare_server
    ├─compare_columns
    ├─%nin%
    ├─info_react_theme
    ├─█─create_new_data
    │ ├─anti_join_base
    │ ├─select_cols
    │ ├─rename_join_col
    │ └─create_new_column
    ├─new_react_theme
    ├─█─create_deleted_data
    │ ├─anti_join_base
    │ ├─select_cols
    │ ├─rename_join_col
    │ └─create_new_column
    ├─deleted_react_theme
    ├─█─create_modified_data
    │ ├─select_cols
    │ ├─█─diff_values_by_key
    │ │ ├─compare_values
    │ │ └─class_diffs
    │ ├─rename_join_col
    │ ├─create_new_column
    │ └─class_diffs
    ├─select_cols
    ├─changed_react_theme
    ├─left_join_base
    └─█─create_comparison_report
      ├─█─create_new_data
      │ ├─anti_join_base
      │ ├─select_cols
      │ ├─rename_join_col
      │ └─create_new_column
      ├─█─create_deleted_data
      │ ├─anti_join_base
      │ ├─select_cols
      │ ├─rename_join_col
      │ └─create_new_column
      ├─█─create_modified_data
      │ ├─select_cols
      │ ├─█─diff_values_by_key
      │ │ ├─compare_values
      │ │ └─class_diffs
      │ ├─rename_join_col
      │ ├─create_new_column
      │ └─class_diffs
      ├─create_empty_tbl
      ├─compare_columns
      ├─class_diffs
      └─█─compare_summary
        └─█─compare_data
          ├─█─create_new_data
          │ ├─anti_join_base
          │ ├─select_cols
          │ ├─rename_join_col
          │ └─create_new_column
          ├─█─create_deleted_data
          │ ├─anti_join_base
          │ ├─select_cols
          │ ├─rename_join_col
          │ └─create_new_column
          ├─█─create_changed_data
          │ ├─select_cols
          │ ├─█─diff_values_by_key
          │ │ ├─compare_values
          │ │ └─class_diffs
          │ ├─rename_join_col
          │ ├─create_new_column
          │ └─class_diffs
          ├─class_diffs
          └─compare_columns

What rows are here now that weren’t here before?

To find new rows, we’ll compare T1Data and T2Data, then check the result against NewData.

T1Data <- dfdiffs::T1Data
T2Data <- dfdiffs::T2Data
NewData <- dfdiffs::NewData

Timepoint 1 data (original)

T1Data represents data collected at the first timepoint (T1).

T1Data: Simulated ‘time-point 1’ data

Timepoint 2 data (new)

T2Data is the ‘new’ dataset, representing data collected at the second timepoint (T2).

T2Data: Simulated ‘time-point 2’ data

create_new_data()

create_new_data() returns the ‘new data’ (i.e., the rows that are here now but weren’t here before). We pass the newer dataset to compare and the original dataset to base:

create_new_data(
  compare = T2Data, 
  base = T1Data)

Output from create_new_data(): Differences between ‘time-point 1’ and ‘time-point 2’

We can confirm this output is correct by checking it against the stored NewData dataset. The two tables should match.

NewData: stored differences from create_new_data()

What rows were here before that aren’t here now?

To test for deleted data, we’ll compare CompleteData and IncompleteData, then check the result against DeletedData.

CompleteData <- dfdiffs::CompleteData
IncompleteData <- dfdiffs::IncompleteData
DeletedData <- dfdiffs::DeletedData

A complete dataset

CompleteData is the full dataset, before any rows were removed.

CompleteData: simulated data for checking ‘deleted data’

An incomplete dataset

IncompleteData is a copy of CompleteData with some rows removed.

IncompleteData: simulated data for checking ‘deleted data’

create_deleted_data()

create_deleted_data() returns the rows in base that are missing from compare. Here, that means the rows in CompleteData that were dropped from IncompleteData:

create_deleted_data(
  compare = IncompleteData, 
  base = CompleteData) 

Output from create_deleted_data(): Differences between CompleteData and IncompleteData

The deleted data

The output above matches the data stored in DeletedData, which confirms the deleted rows were identified correctly.

DeletedData: Output from create_deleted_data()

What values have been changed?

dfdiffs has two functions for this question: create_changed_data() and create_modified_data(). Both build on the same base-R engine (diff_values_by_key() and compare_values()), and they differ only in the shape of their output. create_changed_data() returns $num_diffs and $var_diffs, which compare_data() uses. create_modified_data() returns $diffs_byvar and $diffs, which the app’s compare module (mod_compare) and create_comparison_report() use.

To check for changed values, we’ll compare InitialData and ChangedData:

InitialData <- dfdiffs::InitialData
ChangedData <- dfdiffs::ChangedData

Initial data

InitialData is the original version of the data.

InitialData: simulated data for checking ‘changed/modified data’

Changed data

ChangedData is the version we’ll compare against InitialData.

ChangedData: simulated data for checking ‘modified data’

create_changed_data()

create_changed_data() returns a list of tables. The code below stores the result and prints the names of its elements:

changed <- create_changed_data(
  compare = ChangedData,
  base = InitialData)
names(changed)
#> [1] "num_diffs"   "var_diffs"   "class_diffs"
Counts of changes (num_diffs)

num_diffs stores the number of changes for each variable.

num_diffs: counts of changes for ‘changed data’

Changes by row (var_diffs)

var_diffs lists the changes row by row.

var_diffs: Row-by-row of changes for ‘changed data’

create_modified_data()

create_modified_data() also returns a list of tables, with different element names:

modified <- create_modified_data(
  compare = ChangedData,
  base = InitialData)
names(modified)
#> [1] "diffs"       "diffs_byvar" "class_diffs"
Counts of changes (diffs_byvar)

diffs_byvar stores the number of changes for each variable.

diffs_byvar: counts of changes for ‘modified data’

Changes by row (diffs)

diffs lists the changes row by row.

diffs: Row-by-row changes of ‘modified data’