# install.packages("remotes")
remotes::install_github("andrew-farrey/sysPrep")sysPrep is an R package that provides formalized preprocessing methods
for syndromic surveillance data from the National Syndromic Surveillance
Program (NSSP) Electronic Surveillance System for the Early Notification
of Community-based Epidemics (ESSENCE) API.
Its methods focus on preparing ESSENCE data for real-world, applied
public health surveillance.
sysPrep's functions were developed through troubleshooting repeated
false-positive overdose clusters back to duplicate ESSENCE visit
records: a cause only visible in line-level data pulled via ESSENCE's
dataDetails API endpoint (Full Details). Its functions will not work
on pre-aggregated ESSENCE data pulled using the tableBuilder or
timeSeries API endpoints.
These steps originated in a statewide overdose surveillance system, but
the value they add isn't specific to overdose: sysPrep's methods
broadly improve the quality of downstream ESSENCE data, particularly
wherever ensuring accurate, distinct patient counts is an implicit part
of the surveillance task. Results will vary by site: some sites'
Visit_ID may not support reliable deduplication, and for overdose
surveillance specifically, where a larger share of EMS-responded
overdoses involve refused transport, those patients are never captured
in ED data at all regardless of preprocessing.
These methods are not required to perform case counting, cluster
detection, or anomaly detection with ESSENCE data. Many surveillance
questions tolerate the data quality issues sysPrep addresses without
materially affecting conclusions. They are most valuable for
small-count, high-impact case definitions, where external validity
and minimizing false-positive clusters/anomalies matter most.
Raw ESSENCE data pulls can present four categories of data quality issues that bias case counts or distort cluster/anomaly detection if left unaddressed:
| Problem | Functions |
|---|---|
| Duplicate Records: multiple rows per visit due to visit date changes to the initial record, patient ID corrections, and patient class transitions | dedupe(), summarize_duplicates(), classify_duplicates() |
Non-Emergency Providers: facilities without EDs, and free-standing emergency departments (FSEDs) onboarded to ESSENCE with a non-emergency FacilityType, included in pulls |
filter_care_setting(), review_facility_ed_visits() |
| Mis-Submitted and Invisible Direct Admissions: some inpatient admissions are mistakenly submitted as unrelated to a preceding ED visit (a data quality artifact), and genuine direct admissions are excluded entirely from HasBeenE = 1 queries | link_encounters() |
Out-of-State and OTHER_REGION (unknown residence) Visits: these visits' Region doesn't match any in-state value, so ordinary region-scoped rollups (maps, county summary tables) silently exclude them with no explicit filter required, understating burden at the location where care was actually delivered |
assign_treating_geography(), assign_facility_geography() |
sysPrep synthesizes these methods into a reproducible, documented
pipeline.
sysPrep's functions have been validated against records pulled from the
following NSSP ESSENCE data sources. Both sources return ED visit and
inpatient admission records, so all exported functions apply to either:
ESSENCE datasource code |
Full name | Used by |
|---|---|---|
va_er, va_hosp |
Patient Location (Full Details), Facility Location (Full Details) | dedupe(), summarize_duplicates(), classify_duplicates(), filter_care_setting(), review_facility_ed_visits(), link_encounters(), assign_treating_geography(), assign_facility_geography() |
Column names and expected value formats (e.g., {SITE}_{REGION} region strings,
HasBeenE/HasBeenAdmitted flags) reflect these two data sources. Pulls
from other ESSENCE data sources may require column renaming before use.
library(sysPrep)
# Inspect raw data quality before preprocessing
summarize_duplicates(essence_raw)
classify_duplicates(essence_raw)
# Full preprocessing pipeline
clean <- essence_raw |>
# Step 1: Remove duplicate records (retain most recently transmitted)
dedupe(order_by = Arrived_Date_Time, keep = "last") |>
# Step 2: Filter to ED and inpatient facilities; correct FSED types
filter_care_setting(
fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
) |>
# Step 3: Assign treating facility geography to out-of-state visits
# (writes region_hybrid/zip_code_hybrid; Region/ZipCode stay untouched)
assign_treating_geography()
# Check for facility-level anomalies
clean |> review_facility_ed_visits(method = "both", date_col = Date)For encounter linkage with a separate inpatient pull:
# Deduplicate both pulls
ed_clean <- essence_ed |> dedupe(order_by = Arrived_Date_Time)
inpatient_clean <- essence_inpatient |> dedupe(order_by = Arrived_Date_Time)
# Link into care episodes (captures direct admissions);
# one merged row per true encounter by default
episodes <- link_encounters(ed_clean, inpatient_clean)
# Distribution of care pathways
episodes |>
dplyr::count(.patient_class_sequence, sort = TRUE)For comparing residential vs. treating geography side by side, both
geography functions write to new columns by default, so chaining them adds
region_hybrid and region_facility without disturbing region itself
(built fresh here rather than reusing clean above, since clean already
ran assign_treating_geography() once as step 3 of its own pipeline):
geo_all <- essence_raw |>
dedupe(order_by = Arrived_Date_Time, keep = "last") |>
filter_care_setting(
fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
) |>
assign_treating_geography() |>
assign_facility_geography()
geo_all |>
dplyr::select(hospital_name, region, region_hybrid, region_facility)
# region: untouched original
# region_hybrid: residential geography, with treating geography substituted
# only for out-of-state/unknown-residence visits
# region_facility: treating geography for every visitFull documentation including function reference pages and methodological vignettes is available at: https://andrew-farrey.github.io/sysPrep/
Vignettes:
- Getting Started
- Understanding and Resolving Duplicate Records
- Linking ED and Inpatient Records
- Geographic Attribution
Rnssp: NSSP ESSENCE API access, alerting algorithms, and syndromic surveillance utilities.sysPrepis designed to operate on data returned byRnsspAPI calls; it offers an optional preprocessing layer between raw API output and case counting or cluster detection, for surveillance programs where that layer adds value.
sysPrep relies heavily on several packages whose authors deserve
explicit credit:
-
janitor(Sam Firke):clean_names()is called on entry and exit of every function in the package. The column-name agnosticism that letssysPrepaccept both raw ESSENCE PascalCase and post-clean_names()snake_case is built entirely on this foundation. -
dplyr,rlang, andtidyr(Hadley Wickham, Lionel Henry, and the tidyverse team): the data manipulation backbone and theinform()/warn()/abort()messaging infrastructure used throughout. -
cli(Gábor Csárdi): the rich formatted output for theprint.essence_dup_summary()andprint.essence_dup_classified()S3 methods. -
Rnssp(Gbedegnon Roseric Azondekon, Michael Sheppard, and the CDC BioSense team): the upstream package that handles NSSP ESSENCE API authentication and data retrieval.sysPrepwould have no data to preprocess without it.
The methods formalized in sysPrep were developed through applied drug
overdose surveillance and cluster detection work using NSSP ESSENCE data.
If you use sysPrep in published work, please cite the package directly:
citation("sysPrep")Citation metadata is also available in
CITATION.cff.
MIT © Andrew Farrey
