Off policy evaluation audit for logged bandits: measures how much of a target policy's probability mass sits outside what the log could have produced, shows that a 95 percent interval around the standard estimator covers 0.26 of the time, and returns exit 2 rather than a number when the estimate would be about a different quantity.
python numpy ci-cd recommender-systems causal-inference model-validation importance-sampling contextual-bandits mlops github-actions off-policy-evaluation doubly-robust ml-evaluation counterfactual-evaluation inverse-propensity-scoring
-
Updated
Aug 26, 2026 - Python