Skip to content

Repository files navigation

CausalEstimate

Unittests Lint using flake8 Formatting using black PyPI version Python versions License: MIT Docs Ask DeepWiki

📖 Documentation: kirilklein.github.io/CausalEstimate

CausalEstimate estimates causal effects from observational data using TMLE, AIPW, inverse probability weighting, and matching. Provide the propensity scores and outcome predictions from your own models as columns in a pandas DataFrame.


Why CausalEstimate?

Many causal-inference libraries combine model fitting and effect estimation. CausalEstimate keeps the two separate:

  • Bring your own predictions. Fit propensity and outcome models with scikit-learn, XGBoost, a deep model, or an external system. CausalEstimate uses the resulting columns to estimate effects.
  • Pandas-native. Pass a DataFrame with named columns and get back a plain dictionary.
  • Focused. Estimate ATE, ATT, ATC, and risk ratios. TMLE, AIPW and IPW include influence-curve standard errors, and every estimator supports bootstrap confidence intervals. Built-in diagnostics help assess overlap, covariate balance, and weights.

Choose DoWhy or EconML instead if you need an end-to-end modeling pipeline, causal-graph construction, or heterogeneous treatment effects.


Installation

pip install CausalEstimate

Quickstart

Single estimator

Each estimator is configured with its column names and effect type, then called with compute_effect(df).

from CausalEstimate.datasets import load_binary_with_probas
from CausalEstimate.estimators import IPW

# Synthetic data with known ground truth. Columns "ps", "probas",
# "probas_t0", "probas_t1" stand in for your own model's predictions.
df, params = load_binary_with_probas(
    n_samples=5000,
    random_state=42,
    return_params=True,
)

ipw = IPW(
    effect_type="ATE",
    treatment_col="treatment",
    outcome_col="Y",
    ps_col="ps",
)
results = ipw.compute_effect(df)

print(f"IPW estimated effect: {results['effect']:.4f}")
print(f"True ATE: {params['true_ate']:.4f}")
IPW estimated effect: 0.1599
True ATE: 0.1644

results["effect"] is the estimated effect. effect_1 and effect_0 are the mean potential outcomes under treatment and control.

To compare several estimators or add bootstrap confidence intervals, see Multiple Estimators & Bootstrap.


What's included

Estimator ATE ATT ATC RR RRT ARR
IPW ✓ ✓ ✓ ✓ ✓ ✓
AIPW ✓ ✓ ✓ ✓ ✓ ✓
TMLE ✓ ✓ ✓ ✓ ✓ ✓
Matching ✓* – – – – ✓*

ATE is the average treatment effect, and ATT and ATC are the average treatment effects among treated and control units. RR is the risk ratio, RRT is the risk ratio among treated units, and ARR is the absolute risk reduction. * With a caliper, the matched population is neither the full nor the treated population; interpret accordingly.

  • Diagnostics (CausalEstimate.diagnostics): covariate balance, positivity and overlap metrics, effective sample size, and weight summaries.
  • Common-support filtering (CausalEstimate.filter_common_support): trim to the propensity-score overlap region.
  • Matching (CausalEstimate.estimators.Matching): greedy and optimal propensity-score matching with an optional caliper.
  • Synthetic datasets (CausalEstimate.datasets): load_binary and load_binary_with_probas can return known effects for benchmarking.
  • Plotting (CausalEstimate.vis): propensity-score and outcome-probability distributions by treatment group.

Run all diagnostics in one call before reporting IPW or TMLE estimates:

from CausalEstimate import run_diagnostics

report = run_diagnostics(
    df,
    ps_col="ps",
    treatment_col="treatment",
    covariate_cols=["X1", "X2"],
)
print(report["flags"])  # {"extreme_ps": bool, "unbalanced": bool}
print(report["positivity"], report["weights"], report["balance_summary"])

Documentation

Balance before/after weighting at a glance (notebook):

Love Plot

Propensity-score overlap before and after IPW weighting:

Propensity Score Boxplot

Bootstrap confidence-interval coverage on the built-in synthetic data, where the true effect is known:

Zipper Plot


Contributing

See CONTRIBUTING.md for the dev setup and test workflow. Bug reports and feature requests are welcome as issues; questions can also go to kikl@di.ku.dk.

License

MIT — see LICENSE.

Citation

Use the "Cite this repository" button on GitHub (backed by CITATION.cff), or:

@software{causalestimate,
  author = {Klein, Kiril},
  title = {CausalEstimate: A Python Library for Causal Inference},
  year = {2024},
  url = {https://github.com/kirilklein/CausalEstimate},
  note = {GitHub repository}
}

About

Lightweight causal inference from precomputed propensity scores and outcome predictions: IPW, AIPW, TMLE, matching, bootstrap CIs. Pandas-native.

Topics

Resources

Contributing

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages