A synthetic retail-media experimentation and analytics project for measuring whether advertising drives incremental customer behavior, rather than simply receiving credit for purchases that may have happened anyway.
Retail-media teams routinely report metrics such as impressions, clicks, attributed revenue, and ROAS. These metrics answer an important question:
Which campaign received credit for a purchase?
But they do not answer a different and often more important business question:
Did the advertising actually cause additional purchases or revenue?
A customer who buys after seeing an ad may have purchased even without the campaign. Relying only on attributed revenue can therefore overstate advertising effectiveness.
This project simulates how a retailer or retail-media network can use randomized treatment and control groups to estimate campaign incrementality and make more defensible budget decisions.
The core distinction is:
Attribution: How much revenue is credited to the campaign?
Incrementality: How much additional customer behavior occurred because shoppers were assigned to the advertising treatment?
The project creates a synthetic retail-media environment and implements the analytical workflow from campaign generation through commercial decision support.
Synthetic shoppers, advertisers, and campaigns
|
v
Randomized treatment / control assignment
|
v
Simulated advertising and transactions
|
v
PostgreSQL raw.*
|
v
dbt staging / intermediate / marts
|
v
Treatment-control lift point estimates
|
v
Python uncertainty + experiment health + decisions
|
v
Processed CSVs and reporting views
Airflow orchestrates the warehouse rebuild, Python enrichment, and
reporting-layer build. It does not compute lift, health, iROAS, or decisions.
Tableau, if used, reads reporting views or CSV extracts of those views.
The system combines:
- Python for reproducible synthetic data generation, experimental assignment, database loading, statistical inference, and recommendations
- PostgreSQL as the analytical warehouse
- dbt for governed staging/intermediate/mart transformations, grain documentation, lineage, and data tests
- Randomized A/B testing for treatment-control incrementality measurement
- Commercial analytics for connecting campaign performance to budget recommendations
- pytest for validating data generation, experiment assignment, inference, dbt parity, DAG structure, and reporting-layer copies
- Apache Airflow (optional extra) for local orchestration of that same sequence
- Tableau-ready reporting views (optional BI consumption; the repository does not ship a
.twb)
The experimental unit is a member within a campaign.
Eligible members are selected based on the campaign's retailer and target audience segment.
Members randomized to treatment are eligible to receive campaign advertising.
Members randomized to control are held out from campaign advertising.
The current simulation verifies that:
Control-group ad events = 0
This prevents advertising contamination of the holdout group.
The current seed-0 simulation contains:
| Experiment metric | Value |
|---|---|
| Campaigns | 24 |
| Campaign-member assignments | 60,357 |
| Treatment assignments | 48,295 |
| Control assignments | 12,062 |
| Treatment share | 80.02% |
| Control share | 19.98% |
The assignment procedure targets an approximately 80/20 treatment-control split while maintaining roughly 500 control members per campaign.
The project contains two distinct measurement layers.
Standard retail-media metrics include:
- impressions
- clicks
- CTR
- spend
- CPC
- orders
- conversion rate
- attributed revenue
- ROAS
- average order value
These describe campaign delivery and attributed performance.
For each campaign, customer outcomes are compared between randomized treatment and control groups.
The primary campaign-level estimator is a difference in means:
$$ \hat{\tau}=
\bar{Y}_T - \bar{Y}_C $$
where (Y) can represent conversion, orders per member, or revenue per member.
For conversion:
$$ \text{Absolute Lift}=
\hat{p}_T - \hat{p}_C $$
where (\hat{p}_T) and (\hat{p}_C) are the treatment and control conversion rates.
Estimated incremental revenue is calculated as:
$$ \widehat{\text{Incremental Revenue}}=
n_T \left( \bar{Y}_T^{\text{revenue}}=
\bar{Y}_C^{\text{revenue}} \right) $$
where (n_T) is the number of treatment members and (\bar{Y}_T^{\text{revenue}}) and (\bar{Y}_C^{\text{revenue}}) are average revenue per member in the treatment and control groups.
This is an ITT-style treatment-effect estimate for the synthetic randomized experiment.
All results below are generated from synthetic data. They demonstrate the behavior of the analytical system and should not be interpreted as real retailer or advertiser performance.
| Metric | Current seed-0 result |
|---|---|
| Synthetic members | 25,000 |
| Advertisers | 8 |
| Campaigns | 24 |
| Impressions | 3,910,259 |
| Clicks | 23,993 |
| CTR | 0.614% |
| Ad spend | $46,678.60 |
| Attributed orders | 1,360 |
| Attributed revenue | $57,964.34 |
| ROAS | 1.24 |
| Average revenue per order | $42.62 |
Across the current campaign experiments:
| Metric | Result |
|---|---|
| Overall treatment conversion rate | 5.84% |
| Overall control conversion rate | 3.03% |
| Overall treatment-control difference | +2.81 percentage points |
| Median campaign treatment CVR | 5.08% |
| Median campaign control CVR | 2.43% |
| Median campaign absolute lift | +2.61 percentage points |
| Maximum campaign absolute lift | +5.57 percentage points |
| Estimated incremental orders | ~1,401 |
| Estimated incremental revenue | ~$58,802 |
The strongest current point estimate is Campaign 9:
Treatment CVR: 9.51%
Control CVR: 3.94%
Absolute lift: +5.57 percentage points
Incremental revenue: ~$4,626
These are point estimates from a synthetic randomized experiment, now reported with statistical uncertainty (analytic 95% CIs for conversion lift; member-level bootstrap 95% CIs for incremental orders and revenue). They are not live advertiser results.
The central analytical idea of the project is that attributed performance and incremental performance are not the same quantity.
The current synthetic run reports:
Attributed revenue: $57,964
Experimentally estimated
incremental revenue: $58,802
These values happen to be similar in the current simulation, but they are produced by different measurement approaches.
Attributed revenue assigns campaign credit to transactions.
Incremental revenue estimates what additional revenue occurred relative to the randomized control group's counterfactual baseline.
The purpose of the project is not to demonstrate a particular dollar result. It is to demonstrate the analytical framework required to distinguish the two.
The Stage 4 decision table makes the same distinction with two return metrics:
ROAS = attributed revenue / spend
iROAS = estimated incremental revenue / spend
Attributed ROAS answers how much credited revenue was associated with a dollar of spend. iROAS answers how much experimentally incremental revenue was estimated per dollar of spend. They can disagree.
Before a treatment-effect estimate is used for a budget decision, the project checks whether the randomized experiment is structurally trustworthy.
Campaign-level diagnostics include:
- randomized-arm counts versus the campaign's stored holdout fraction
- sample-ratio mismatch (chi-square goodness-of-fit)
- control-group advertising leakage
- duplicate
campaign_id × member_idassignments - assignment-to-outcome completeness
- pre-treatment standardized differences for signup tenure and pre-campaign conversion
Overall status is one of PASS, WARN, or FAIL.
FAIL is reserved for structural problems that make causal interpretation unsafe: control exposure leakage, missing assigned members in the outcome table, duplicate randomized assignments, or severe SRM (p < 0.001). Mild SRM (0.001 ≤ p < 0.05) or a pre-treatment standardized difference above 0.10 produces WARN and does not automatically invalidate the experiment.
Planned power and planned MDE are not reconstructed. The simulator never stored a pre-registered design calculation, so those metadata fields are left null. After the experiment, precision is summarized with confidence-interval width rather than post-hoc observed power.
For the current seed-0 run:
PASS: 22 campaigns
WARN: 2 campaigns
FAIL: 0 campaigns
Control impressions on holdout members: 0
Missing assigned outcomes: 0
There are two deterministic outputs.
data/processed/campaign_recommendations.csv maps incrementality efficiency flags to a coarse action. It does not use experiment health or interval estimates.
| Efficiency flag | Recommendation |
|---|---|
high_impact |
increase_budget |
moderate |
maintain |
low_impact |
monitor |
| inefficient campaign | reduce_budget |
Current seed-0 mapping: 2 increase / 10 maintain / 12 monitor / 0 reduce.
data/processed/campaign_measurement_decisions.csv is the Stage 4 decision table. It combines attributed ROAS, incremental revenue and iROAS with uncertainty, experiment health, and efficiency flags.
Decision order is:
- If experiment health is
FAIL, assigndo_not_interpret. - Otherwise use the incremental-revenue 95% CI, not a p-value cutoff.
- If that CI is entirely positive and iROAS ≥ 1, assign
increase_budget. - If that CI is entirely positive but iROAS < 1, assign
maintain. - If that CI is entirely negative, assign
reduce_budget. - If the CI includes zero, assign
inconclusive.
A significant-looking point estimate is not enough. A wide interval that includes zero is treated as inconclusive. A failed experiment is not interpreted even if the point estimate is large and positive.
Current seed-0 health-aware decisions:
| Decision | Campaigns |
|---|---|
| Increase budget | 14 |
| Maintain | 3 |
| Inconclusive | 7 |
| Reduce budget | 0 |
| Do not interpret | 0 |
Both layers are rule based. Neither is a trained recommendation model.
PostgreSQL remains the warehouse. dbt is the canonical transformation layer for the measurement path used by incrementality, ROAS/iROAS, and campaign decisions.
dbt was introduced so experiment, attribution, and campaign-measurement inputs can be rebuilt from documented, tested models rather than a loose collection of SQL scripts.
This remains a synthetic portfolio system, not a production warehouse or cloud data platform.
dbt owns:
- source declarations over
raw.* - staging type cleanup
- intermediate assigned populations
- incrementality and attribution marts
- model grains, documentation, and lineage
- generic and experiment-integrity tests
Python still owns:
- Wald conversion inference
- member-level bootstrap intervals
- SRM p-values and baseline SMDs
- iROAS (incremental revenue / spend)
- PASS/WARN/FAIL health status
- deterministic campaign decisions
Airflow owns:
- task order
- retries for transient failures
- failing the pipeline when dbt tests or publication checks fail
Frozen v1 estimands and Stage 4 decision rules were not rewritten in SQL. Parity tests compare dbt marts to the existing pandas ITT reference.
Incrementality (randomized assignment + campaign-window purchases; ignores source_campaign_id):
raw.campaign_experiment_assignments
↓
staging.stg_experiment_assignment
↓
intermediate.int_experiment_assigned_population
↓
marts.experiment_member_outcomes
↓
marts.experiment_lift_metrics (point estimates)
marts.segment_performance_metrics
marts.experiment_health_metrics (integrity inputs)
↓
Python inference, health status, iROAS, decisions
Attribution (credit via source_campaign_id):
staging.stg_ad_events
staging.stg_transactions.source_campaign_id
↓
marts.campaign_base_metrics
↓
marts.campaign_spend_metrics (attributed ROAS)
These branches remain separate. Attributed ROAS is not iROAS.
raw untyped CSV loads
staging typed dbt staging models
intermediate assigned campaign-member population
marts business marts, including Python-enriched tables
reporting BI views over trusted marts (Stage 7; built after Python)
Airflow wraps this path. It does not replace dbt or Python:
Airflow DAG
|
+--------------------+--------------------+
| | |
v v v
check sources dbt run/test Python inference
(excl. reporting) |
v
health + decisions
|
v
reporting views + extracts
|
v
processed CSV checks
Files under sql/staging/ and sql/marts/ are frozen legacy references. They are not the default rebuild path. Reporting marts not migrated in Stage 5 still live only as legacy SQL:
sql/marts/campaign_funnel_metrics.sql
sql/marts/daily_campaign_trends.sql
sql/marts/executive_summary_metrics.sql
sql/marts/campaign_incrementality_rankings.sql
marts.campaign_measurement_decisions is created by Python, not by dbt.
retail-media-platform/
├── app/
│ ├── core/
│ │ └── database.py
│ ├── orchestration/
│ │ └── checks.py
│ └── statistics/
│ ├── experiment_inference.py
│ ├── experiment_health.py
│ └── experiment_decisions.py
│
├── configs/
│ ├── experiment_config.yaml
│ ├── recommendation_config.yaml
│ └── simulation_config.yaml
│
├── data/
│ ├── synthetic/
│ │ ├── members.csv
│ │ ├── advertisers.csv
│ │ ├── campaigns.csv
│ │ ├── campaign_experiment_assignments.csv
│ │ ├── ad_events.csv
│ │ └── transactions.csv
│ │
│ └── processed/
│ ├── experiment_lift_metrics.csv
│ ├── segment_performance_metrics.csv
│ ├── experiment_design_metadata.csv
│ ├── experiment_health_metrics.csv
│ ├── campaign_measurement_decisions.csv
│ ├── campaign_recommendations.csv
│ ├── campaign_decision_dashboard.csv
│ ├── campaign_subgroup_dashboard.csv
│ └── portfolio_decision_summary.csv
│
├── dbt/
│ ├── dbt_project.yml
│ ├── profiles.yml
│ ├── models/
│ ├── macros/
│ └── tests/
│
├── dags/
│ └── retail_media_measurement.py
│
├── docs/
│ ├── attribution_methodology.md
│ ├── experiment_design.md
│ ├── kpi_definitions.md
│ ├── airflow.md
│ ├── tableau.md
│ └── tableau_dashboard_spec.md
│
├── notebooks/
│ ├── 01_data_checks.ipynb
│ ├── 03_segment_lift_analysis.ipynb
│ ├── 04_business_insights.ipynb
│ └── 05_experiment_health_decisions.ipynb
│
├── scripts/
│ ├── generate_members.py
│ ├── generate_advertisers.py
│ ├── generate_campaigns.py
│ ├── assign_experiments.py
│ ├── generate_ad_events.py
│ ├── generate_transactions.py
│ ├── load_to_postgres.py
│ ├── run_dbt.sh
│ ├── run_analytics.sh
│ ├── run_incrementality.py
│ ├── run_experiment_decisions.py
│ ├── run_reporting.sh
│ ├── export_reporting.py
│ ├── check_pipeline_readiness.py
│ ├── validate_pipeline_outputs.py
│ ├── generate_recommendations.py
│ └── validate_dbt_parity.py
│
├── sql/
│ ├── staging/
│ └── marts/
│
├── tests/
├── .env.example
├── requirements.txt
├── requirements-airflow.txt
└── README.md
- pandas
- NumPy
- SQLAlchemy
- psycopg2
- PyYAML / YAML configuration
- PostgreSQL
- SQL
- dbt (dbt-postgres)
- pytest
- dbt tests
- Jupyter notebooks
- Apache Airflow 3.x (optional;
requirements-airflow.txt)
- Tableau Desktop or Tableau Public (external; workbook is assembled from
docs/tableau_dashboard_spec.md) - PostgreSQL
reportingviews and CSV extracts indata/processed/
The current project does not contain a trained machine-learning model, production API, Tableau Server deployment, or MLflow experiment-tracking workflow. Airflow here is a local DAG over frozen synthetic data, not a cloud-native production orchestrator. Tableau is a self-service consumption layer over trusted outputs, not a second measurement engine.
python -m venv venv_rmp
source venv_rmp/bin/activate
pip install -r requirements.txtOptional local orchestrator (after the analytics install):
pip install -r requirements-airflow.txtCreate .env from .env.example and configure the PostgreSQL connection.
Run the generators in dependency order:
python scripts/generate_members.py
python scripts/generate_advertisers.py
python scripts/generate_campaigns.py
python scripts/assign_experiments.py
python scripts/generate_ad_events.py
python scripts/generate_transactions.pyGenerated files are written to:
data/synthetic/
The simulation seed is configured in:
configs/simulation_config.yaml
python scripts/load_to_postgres.pyThis loads the generated CSV files into the PostgreSQL raw schema.
dbt reads POSTGRES_* from .env. The wrapper loads that file:
scripts/run_dbt.sh debug
scripts/run_dbt.sh run
scripts/run_dbt.sh testThis builds typed staging tables, the assigned-population intermediate model, incrementality marts, and the attribution spend/ROAS marts used by iROAS.
Lineage and model docs:
scripts/run_dbt.sh docs generate
scripts/run_dbt.sh docs serveDo not commit generated dbt/target/ artifacts.
Optional convenience command (measurement dbt run/test excluding reporting models, Python enrichment, reporting views, pytest):
scripts/run_analytics.shThat script is the manual/CI equivalent of the Airflow DAG. The DAG is the canonical orchestrated path; the script remains for debugging. The first dbt run uses --exclude tag:reporting so Python can enrich marts before reporting views join them.
dbt produces point estimates and health inputs. Python adds intervals, SRM/SMD status, iROAS, and decisions. Re-running dbt drops those enrichment columns, so Python must run after dbt:
python scripts/run_incrementality.py --export-csv
python scripts/run_experiment_decisions.py --export-csvThis does not regenerate synthetic source data and does not recompute the frozen v1 treatment-effect point estimates.
--legacy-sql on those scripts executes the frozen files under sql/ instead of using dbt output. Prefer dbt run.
After decisions exist:
scripts/run_reporting.shThis builds PostgreSQL views in schema reporting, tests them, and exports Tableau extracts. It does not recompute lift, health, or decisions.
Connect Tableau to reporting.* (local) or to the exported CSVs (Tableau Public). Specification: docs/tableau.md and docs/tableau_dashboard_spec.md. There is no checked-in .twb.
The simple efficiency-flag mapping is unchanged:
python scripts/generate_recommendations.pyOutput:
data/processed/campaign_recommendations.csv
The health-aware decision table is written separately by run_experiment_decisions.py:
data/processed/campaign_measurement_decisions.csv
python -m pytest -qCurrent result after Stage 7:
python -m pytest -q → 148 passed
dbt test --exclude tag:reporting → 107 passed
dbt test --select tag:reporting → 26 passed
DAG import tests require pip install -r requirements-airflow.txt. Analytics tests do not.
See Automated Measurement Pipeline and docs/airflow.md.
Added an Airflow-orchestrated analytical pipeline that coordinates dbt transformations, data-quality validation, randomized incrementality inference, experiment-health evaluation, campaign decision outputs, and the Stage 7 reporting layer.
This is a local portfolio DAG, not a production or cloud orchestration system. The underlying dataset is frozen synthetic seed-0 data. Scheduling is therefore manual (schedule=None, catchup=False). A daily cron would only re-run the same snapshot.
The measurement path was already reproducible, but it required a human to run scripts in the correct order. Stage 6 makes that order explicit, retryable, and failure-safe.
DAG id: retail_media_measurement_pipeline (dags/retail_media_measurement.py)
check_environment
↓
check_database
↓
validate_sources
↓
run_dbt_build (scripts/run_dbt.sh run --exclude tag:reporting)
↓
run_dbt_tests (scripts/run_dbt.sh test --exclude tag:reporting; no retries)
↓
run_incrementality (scripts/run_incrementality.py --export-csv)
↓
run_experiment_decisions
↓
build_reporting_layer (scripts/run_reporting.sh; no retries)
↓
validate_outputs
Airflow does not compute Wald intervals, bootstrap draws, SRM, iROAS, or budget rules. Those stay in app/statistics/. dbt still owns warehouse SQL and tests. Tableau is not published by Airflow.
Dependencies are linear. If dbt tests fail, incrementality and decisions do not run. If incrementality fails, decisions are not published. If reporting-layer tests fail, publication does not run. If the publication check fails, the DAG fails.
Transient tasks (database ping, dbt run, Python scripts) retry twice with a 30-second delay. Quality gates (check_environment, run_dbt_tests, build_reporting_layer, validate_outputs) retry zero times so a failed test is not hidden.
Rerunning the DAG against the same frozen inputs overwrites rather than appends:
- dbt models are tables rebuilt in place
- Python enrichment uses
to_sql(..., if_exists="replace") - CSV exports overwrite files in
data/processed/
Frozen statistical results should not change. The DAG does not regenerate synthetic source data.
export RETAIL_MEDIA_REPO="$(pwd)"
export AIRFLOW_HOME="$RETAIL_MEDIA_REPO/.airflow"
export AIRFLOW__CORE__DAGS_FOLDER="$RETAIL_MEDIA_REPO/dags"
export AIRFLOW__CORE__LOAD_EXAMPLES=false
export AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS=admin:admin
export PYTHONPATH="$RETAIL_MEDIA_REPO${PYTHONPATH:+:$PYTHONPATH}"
airflow db migrate
airflow standaloneTrigger:
airflow dags unpause retail_media_measurement_pipeline
airflow dags trigger retail_media_measurement_pipelineFull developer notes: docs/airflow.md.
Stage 7 exposes the frozen measurement system to product, marketing, and business stakeholders without opening SQL, Python, dbt, or Airflow.
The BI layer is Tableau-ready reporting views (schema reporting) plus CSV extracts. The workbook docs/tableau/retail_media_decision_center.twb visualizes those trusted fields; it does not calculate lift, health, iROAS, or decisions. Remaining dashboards follow docs/tableau_dashboard_spec.md. Connection options: docs/tableau.md.
The Executive Campaign Overview is the first finished dashboard in that Tableau Self-Service Decision Center. It summarizes portfolio spend, attributed revenue and attributed ROAS, estimated incremental revenue and iROAS, experiment health, and campaign decisions. Advertiser, decision, and experiment-health filters make the view interactive. The figures are synthetic seed-0 outputs and do not represent real advertiser or retailer performance.
A stakeholder who needs to see campaign performance, experiment validity, incrementality, attribution-versus-incrementality disagreement, uncertainty, and the recommended action.
- Executive Campaign Overview — spend, attributed ROAS, incremental revenue, portfolio iROAS as a ratio of sums, health and decision counts, priority table.
- Experiment Center — integrity first, then conversion lift, incremental orders/revenue, iROAS, each with 95% CIs, then the trusted decision.
- Attribution vs Incrementality — attributed ROAS versus iROAS. Campaign 24 is the synthetic example where credited ROAS is above 1 while iROAS is below 1 and the decision is
maintain. - Segment / Geo Diagnostics — descriptive randomized subgroup comparisons. Not CATE or uplift. No subgroup confidence intervals.
Tableau consumes measurement_decision, experiment_health_status, ROAS, iROAS, lift, and intervals as already-computed fields. It does not calculate experiment decisions.
Portfolio iROAS is sum(estimated incremental revenue) / sum(spend), not an average of campaign iROAS.
Where the analytical layer already has intervals, the spec requires showing the point and the 95% CI (and a zero or 1.0 reference line). Subgroup pages must stay labeled as descriptive and must not draw invented intervals.
Every dashboard should state that the data are synthetic seed-0 outputs, not real advertiser impact.
No last-successful-run timestamp is stored. Use snapshot_label (synthetic seed-0 frozen snapshot).
One important experiment-integrity check confirms that control members receive no campaign advertising:
SELECT COUNT(*) AS control_ad_events
FROM staging.stg_ad_events ae
JOIN staging.stg_experiment_assignment ea
ON ae.campaign_id = ea.campaign_id
AND ae.member_id = ea.member_id
WHERE ea.experiment_arm = 'control';Expected result:
0
Campaign lift results can be inspected with:
SELECT
campaign_id,
control_member_count,
treatment_member_count,
control_conversion_rate,
treatment_conversion_rate,
absolute_lift,
incremental_orders,
incremental_revenue
FROM marts.experiment_lift_metrics
ORDER BY incremental_revenue DESC;The project tests core simulation and analytical logic, including:
- generated data structure
- randomized experiment assignment
- treatment/control logic
- metric calculations
- incrementality calculations
- conversion-lift standard errors, confidence intervals, and p-values
- member-level bootstrap behavior for orders and revenue
- sample-ratio mismatch, control leakage, and outcome-completeness diagnostics
- iROAS transformation from incremental revenue and spend
- health-aware deterministic decisions, including withholding interpretation after a failed experiment
- dbt mart parity against the frozen pandas ITT reference
- warehouse tests for control-exposure isolation, assignment uniqueness, outcome completeness, and subgroup reconciliation
- pipeline readiness and processed-output publication checks
- Airflow DAG import, task graph, retries, catchup/schedule, and reporting-layer task order
- reporting views copy trusted ROAS, iROAS, incremental revenue, health, and decisions without recalculation
- subgroup reporting grain uniqueness and reconciliation to campaign totals
- portfolio iROAS is a ratio of sums, not a mean of campaign iROAS
This project is intentionally focused on randomized retail-media measurement rather than broad platform development.
When an arm has conversion rate 0 or 1, the binomial variance estimate is zero. The implementation returns a collapsed interval rather than dividing by zero; that is a limitation of the Wald SE, not infinite precision.
Orders and revenue use a stratified member-level percentile bootstrap. Bias-corrected intervals are not implemented.
All customers, advertising activity, transactions, campaign results, and revenue are simulated. Reported dollar values demonstrate the analytical workflow and are not real business outcomes.
Within a campaign, eligibility is retailer plus target audience segment, so segment does not vary. Geography is sparse at the control-arm cell size. The implemented checks are standardized differences on signup tenure and pre-campaign conversion. A large standardized difference produces WARN only.
Segment-level analysis is a descriptive subgroup analysis of the same campaign-window RCT outcomes as the primary experiment, not CATE or individualized treatment-effect modeling.
The primary causal design is randomized treatment/control assignment.
The repository currently focuses on Python, PostgreSQL, experimentation, analytical decision support, a local Airflow DAG, and Tableau-ready reporting views rather than a deployed Tableau Server, production API, or cloud dashboard.
Stage 7 ships tested reporting models, extracts, and a dashboard specification. It does not ship a .twb / .twbx. Connecting Tableau and laying out sheets is a manual Desktop/Public step.
The snapshot is frozen synthetic seed 0. There is no trustworthy last-successful-Airflow timestamp to display.
The DAG is not a production orchestrator. It does not ingest live campaigns, run on a cluster, or imply real-time measurement. schedule=None because the synthetic snapshot does not arrive daily.
The simulator did not store a pre-registered minimum detectable effect or power calculation. Those metadata fields are null by design. Interval width is used as the post-experiment precision summary.
The highest-value remaining additions are:
- Assemble and optionally publish the Tableau Public workbook from
docs/tableau_dashboard_spec.mdusing the synthetic CSV extracts. - Optionally migrate remaining reporting marts (
campaign_funnel_metrics,daily_campaign_trends,executive_summary_metrics,campaign_incrementality_rankings) into dbt.
Retail-media measurement is not just a reporting problem.
A campaign can receive credit for customer purchases without causing those purchases. Reliable media decisions therefore require a counterfactual:
What would these customers have done without the advertising?
This project demonstrates how randomized holdouts, governed SQL transformations, treatment-effect estimation, experiment-integrity checks, deterministic decision rules, explicit pipeline orchestration, and a Tableau-ready self-service reporting layer can turn that question into measurable campaign outcomes and interpretable business decisions.
