getting-started
getting-started.RmdMotivation
When a new version of a dataset arrives, the first thing we usually
want to know is what changed since the last one. dfdiffs
answers that with three questions:
- What rows are here now that weren’t here before?
- What rows were here before that aren’t here now?
- What values have been changed?
Packages
We’ll start by loading dfdiffs:
The code below loads the other packages used in this vignette:
Package functions
dfdiffs has a function for each question above, and each
function comes with a pair of example datasets that show how it works.
We’ll walk through each pair below.
Call structure
The four comparison functions are built from a few small helpers, and
the same functions power the package’s Shiny app. I generated the call
trees in this section from the package source with stackcallr
(pak::pak("mjfrigaard/stackcallr")), and they only show
functions defined in dfdiffs.
library(stackcallr)
call_tree_dir("R", root = "create_new_data")
call_tree_dir("R", root = "create_deleted_data")
call_tree_dir("R", root = "create_changed_data")
call_tree_dir("R", root = "create_modified_data")create_new_data() and create_deleted_data()
answer the first two questions (new rows and deleted rows), and both
rely on the same four helpers: anti_join_base(),
select_cols(), rename_join_col(), and
create_new_column(). These helpers are written in base R,
so no external comparison package is involved:
█─create_new_data
├─anti_join_base
├─select_cols
├─rename_join_col
└─create_new_column
█─create_deleted_data
├─anti_join_base
├─select_cols
├─rename_join_col
└─create_new_column
create_changed_data() and
create_modified_data() answer the third question (changed
values), and both build on diff_values_by_key().
diff_values_by_key() compares matched rows column by column
with compare_values(), which is tolerance-aware for numeric
and date-time values and label-aware for factors. Any class mismatches
are recorded with class_diffs(). The call tree for
create_changed_data() is shown below:
█─create_changed_data
├─select_cols
├─█─diff_values_by_key
│ ├─compare_values
│ └─class_diffs
├─rename_join_col
├─create_new_column
└─class_diffs
The Shiny app (launch_app()) wires these functions
together with three modules (upload, select, and compare), and each
module has a UI function and a server function. app_ui()
and app_server() assemble the modules, and
dfdiffs_fresh_theme() supplies the app’s theme. The code
below generates the full call tree for the app:
call_tree_dir("R", root = "launch_app")█─launch_app
├─█─app_ui
│ ├─dfdiffs_fresh_theme
│ ├─mod_upload_ui
│ ├─mod_select_ui
│ └─mod_compare_ui
└─█─app_server
├─█─mod_upload_server
│ ├─█─upload_data
│ │ └─load_flat_file
│ ├─base_react_theme
│ └─comp_react_theme
├─█─mod_select_server
│ ├─select_cols
│ ├─base_react_theme
│ ├─comp_react_theme
│ ├─info_react_theme
│ └─█─create_join_column
│ └─select_cols
└─█─mod_compare_server
├─compare_columns
├─%nin%
├─info_react_theme
├─█─create_new_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─new_react_theme
├─█─create_deleted_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─deleted_react_theme
├─█─create_modified_data
│ ├─select_cols
│ ├─█─diff_values_by_key
│ │ ├─compare_values
│ │ └─class_diffs
│ ├─rename_join_col
│ ├─create_new_column
│ └─class_diffs
├─select_cols
├─changed_react_theme
├─left_join_base
└─█─create_comparison_report
├─█─create_new_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─█─create_deleted_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─█─create_modified_data
│ ├─select_cols
│ ├─█─diff_values_by_key
│ │ ├─compare_values
│ │ └─class_diffs
│ ├─rename_join_col
│ ├─create_new_column
│ └─class_diffs
├─create_empty_tbl
├─compare_columns
├─class_diffs
└─█─compare_summary
└─█─compare_data
├─█─create_new_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─█─create_deleted_data
│ ├─anti_join_base
│ ├─select_cols
│ ├─rename_join_col
│ └─create_new_column
├─█─create_changed_data
│ ├─select_cols
│ ├─█─diff_values_by_key
│ │ ├─compare_values
│ │ └─class_diffs
│ ├─rename_join_col
│ ├─create_new_column
│ └─class_diffs
├─class_diffs
└─compare_columns
What rows are here now that weren’t here before?
To find new rows, we’ll compare T1Data and
T2Data, then check the result against
NewData.
Timepoint 1 data (original)
T1Data represents data collected at the first timepoint
(T1).
T1Data: Simulated ‘time-point 1’
data
Timepoint 2 data (new)
T2Data is the ‘new’ dataset, representing data collected
at the second timepoint (T2).
T2Data: Simulated ‘time-point 2’
data
create_new_data()
create_new_data() returns the ‘new data’ (i.e., the rows
that are here now but weren’t here before). We pass the newer dataset to
compare and the original dataset to base:
create_new_data(
compare = T2Data,
base = T1Data)Output from create_new_data():
Differences between ‘time-point 1’ and ‘time-point 2’
We can confirm this output is correct by checking it against the
stored NewData dataset. The two tables should match.
NewData: stored differences from
create_new_data()
What rows were here before that aren’t here now?
To test for deleted data, we’ll compare CompleteData and
IncompleteData, then check the result against
DeletedData.
CompleteData <- dfdiffs::CompleteData
IncompleteData <- dfdiffs::IncompleteData
DeletedData <- dfdiffs::DeletedDataA complete dataset
CompleteData is the full dataset, before any rows were
removed.
CompleteData: simulated data for
checking ‘deleted data’
An incomplete dataset
IncompleteData is a copy of CompleteData
with some rows removed.
IncompleteData: simulated data for
checking ‘deleted data’
create_deleted_data()
create_deleted_data() returns the rows in
base that are missing from compare. Here, that
means the rows in CompleteData that were dropped from
IncompleteData:
create_deleted_data(
compare = IncompleteData,
base = CompleteData) Output from create_deleted_data():
Differences between CompleteData and
IncompleteData
The deleted data
The output above matches the data stored in DeletedData,
which confirms the deleted rows were identified correctly.
DeletedData: Output from
create_deleted_data()
What values have been changed?
dfdiffs has two functions for this question:
create_changed_data() and
create_modified_data(). Both build on the same base-R
engine (diff_values_by_key() and
compare_values()), and they differ only in the shape of
their output. create_changed_data() returns
$num_diffs and $var_diffs, which
compare_data() uses. create_modified_data()
returns $diffs_byvar and $diffs, which the
app’s compare module (mod_compare) and
create_comparison_report() use.
To check for changed values, we’ll compare InitialData
and ChangedData:
InitialData <- dfdiffs::InitialData
ChangedData <- dfdiffs::ChangedDataInitial data
InitialData is the original version of the data.
InitialData: simulated data for
checking ‘changed/modified data’
Changed data
ChangedData is the version we’ll compare against
InitialData.
ChangedData: simulated data for
checking ‘modified data’
create_changed_data()
create_changed_data() returns a list of tables. The code
below stores the result and prints the names of its elements:
changed <- create_changed_data(
compare = ChangedData,
base = InitialData)
names(changed)
#> [1] "num_diffs" "var_diffs" "class_diffs"
create_modified_data()
create_modified_data() also returns a list of tables,
with different element names:
modified <- create_modified_data(
compare = ChangedData,
base = InitialData)
names(modified)
#> [1] "diffs" "diffs_byvar" "class_diffs"