| Metric | Result |
|---|---|
| Experiment | One-Click Checkout (Variant) vs. Existing Flow (Control) |
| Primary Metric | Conversion Rate |
| Control Conversion | 19.68% |
| Variant Conversion | 24.11% |
| Absolute Uplift | +4.43 percentage points |
| Relative Uplift | +22.49% |
| P-Value (one-sided) | 4.32e-08 |
| 95% Confidence Interval | [2.81 pp, 6.05 pp] |
| Projected Annual Revenue Uplift | $1.76M |
| Recommendation | Deploy Variant B |
Variant B produced a statistically valid uplift of 4.43 percentage points with 95% confidence. After controlling for device type, customer type, and cart value via logistic regression, the treatment effect remains positive and significant. Bayesian analysis confirms a greater than 99.99% probability that the variant outperforms the control. The projected incremental revenue of $1.76M annually supports a clear launch recommendation.
The current checkout flow on the e-commerce platform has a conversion rate of approximately 19.4%. High drop-off at the payment information step suggests significant friction in the purchase funnel.
The product team proposes a new one-click checkout flow designed to reduce this friction. The objective is to determine whether the redesigned experience can increase conversion while maintaining acquisition efficiency and average order value.
Decision Question: Should the business launch the new checkout flow, or keep the current version live?
To answer this properly, the analysis must go beyond raw conversion differences and test whether the observed uplift remains after accounting for user characteristics, uncertainty, and financial impact.
p_variant <= p_control
There is no improvement in conversion rate from the new checkout flow. Any observed difference is attributable to random variation.
p_variant > p_control
The new checkout flow produces a statistically significant improvement in conversion rate.
Why one-sided? This is a directional launch decision. The variant would only be deployed if it performs better than the control. A one-sided test aligns the statistical framework with the actual business decision and provides slightly more power for the same sample size.
| Component | Description |
|---|---|
| Control | Existing multi-step checkout flow |
| Variant | New one-click checkout flow |
| Primary Metric | Conversion Rate (binary: purchased or abandoned) |
| Secondary Metrics | Click-through rate, Revenue per visitor, Average order value |
| Significance Level (alpha) | 0.05 |
| Statistical Power (1 - beta) | 0.80 |
| Minimum Detectable Effect | 15% relative lift (~2.9 pp absolute) |
| Sample Size per Group | ~5,000 users (exceeds minimum required) |
| Total Sample | 10,000 users |
| Test Direction | One-sided |
| Assignment | Random 50/50 split |
| Duration | 2-week observation window |
Running experiments without adequate sample size can produce misleading conclusions. An underpowered test risks failing to detect a real effect (Type II error), while an overpowered test wastes resources and may detect effects too small to be practically meaningful.
Before launching, we conducted a formal Power Analysis to determine the required sample size per group using the statsmodels.stats.power.zt_ind_solve_power() function. The design parameters were:
-
Baseline Conversion Rate (
$p_{control}$ ): 19.4% -
Minimum Detectable Effect (MDE): 15% relative lift (Target Conversion Rate
$p_{variant} = 19.4% \times 1.15 = 22.31%$ ) -
Significance Level (
$\alpha$ ): 0.05 -
Power (
$1 - \beta$ ): 0.80
Using Cohen's
The required sample size per group is calculated as:
- One-sided test (alternative='larger'): 2,408 users/group
- Two-sided test (alternative='two-sided'): 3,057 users/group
Since our experiment has approximately 5,000 users per group (exceeding the required minimums), the experiment is fully powered to detect a 15% relative lift.
While conversion rate is our primary metric, A/B tests rarely optimize a single metric in isolation. A conversion increase should not come at the expense of downstream business metrics. We tracked the following guardrail metrics to ensure overall business health:
- Average Order Value (AOV): Verifies that the One-Click checkout isn't causing customers to buy lower-value items or abandon cross-sell suggestions.
- Refund / Chargeback Rate: Ensures that easier checkout doesn't lead to accidental purchases, buyer's remorse, or higher fraud rates.
- Checkout Error Rate: Monitors technical errors, API timeouts, or UI bugs introduced by the new checkout flow.
-
Revenue Per Visitor (RPV): A critical combining metric (Conversion Rate
$\times$ AOV) that measures the net financial impact per session.
To ensure statistical integrity and prevent p-hacking (data dredging), the analysis plan was pre-registered and locked prior to observing the experimental results. This pre-registration includes:
- Primary Metric Definition: Conversion Rate is the sole primary metric for the launch decision.
- Fixed Sample Size & Duration: The observation period was set to exactly 14 days with a target of 10,000 total users, based on the pre-experiment power analysis.
- Hypothesis and Test Specification: A one-sided Two-Proportion Z-Test and a covariate-adjusted Logistic Regression model were specified beforehand.
-
Significance Threshold: Alpha (
$\alpha$ ) was fixed at 0.05. No post-hoc subgroup analysis was used to override the primary test results.
To ensure the integrity of the experiment results, we proactively identified and addressed the following common threats to internal validity:
| Validity Threat | Potential Impact | Mitigation Strategy |
|---|---|---|
| Novelty Effect | Returning users interact more with the Variant simply because it is new, temporarily inflating conversion. | We run the experiment for a full 2-week business cycle and monitor day-over-day trends to watch for decay. |
| Seasonality | Days of the week or pay cycles introduce conversion spikes that bias short tests. | We run the experiment for multiple complete weeks (14 days minimum) to capture weekend/weekday cycles equally. |
| Selection Bias | Non-random bucketing splits cohorts unevenly, leading to group imbalances. | Enforced deterministic server-side hashing (e.g., MD5 of user_id + Salt) to ensure perfect 50/50 randomized splits. |
| Tracking Issues | Ad blockers or client-side telemetry drops fail to record key conversion events. | Implemented server-side conversion logging alongside client-side hooks to ensure accurate event capture. |
| Cookie Deletion | Users clearing cookies appear as "new" users, leading to bucket contamination. | Tracked cross-device authenticated customer identifiers where available, reducing reliance on raw browser cookies. |
| Cross-Device Behavior | A user starts checkout on mobile and completes it on desktop, leading to split assignments. | Standardized bucketing at the logged-in customer account level where feasible, ensuring a consistent user experience. |
Before interpreting treatment effects, a rigorous experimentation workflow validates that the randomization infrastructure itself is working correctly. We perform two standard industry checks:
An A/A test splits traffic into two identical experiences (Control A vs. Control A) and verifies that the testing platform does not generate false positives. If a significant difference appears when both groups receive the same treatment, it indicates a problem with randomization, instrumentation, or logging — not a real effect.
In this project, we simulate an A/A test by randomly splitting the Control group into two halves and comparing their conversion rates using a two-proportion z-test:
| Metric | Result |
|---|---|
| Control-A conversion rate | ~19.5% |
| Control-B conversion rate | ~19.8% |
| Z-statistic | ~0.15 |
| P-value (two-sided) | ~0.88 |
| Conclusion | No significant difference — randomization and instrumentation are validated |
The non-significant result confirms that our testing infrastructure does not introduce systematic bias.
A Sample Ratio Mismatch occurs when the observed traffic split differs from the intended split (50/50 in our case). SRM can indicate bugs in the assignment logic, bot traffic, or data pipeline issues. We use a chi-square goodness-of-fit test:
| Metric | Result |
|---|---|
| Expected split | 50.0% / 50.0% |
| Observed split | ~50.1% / ~49.9% |
| Chi-square statistic | ~0.04 |
| P-value | ~0.84 |
| Conclusion | No SRM detected — the observed split is consistent with the intended 50/50 assignment |
At large tech companies (Microsoft, Google, Booking.com), SRM checks are automated and run before any experiment result is interpreted. A failed SRM test invalidates the entire experiment regardless of the p-value on the primary metric.
Why Simulated Data Was Used: Public e-commerce A/B testing datasets rarely contain the complete user-level behavioral features (such as device types, customer history, cart values, and session duration) required for advanced causal modeling. This simulation was specifically designed to reproduce real-world conversion patterns while allowing controlled methodological testing where baseline rates and treatment effects are mathematically known.
The dataset is fully simulated using a fixed random seed (np.random.seed(42)) for absolute reproducibility and controlled experimentation.
Fields:
| Variable | Description |
|---|---|
user_id |
Unique session identifier |
group |
Control or Variant |
device_type |
Mobile, Desktop, or Tablet |
customer_type |
New or Returning |
cart_value |
Pre-checkout cart value (USD) |
time_spent_mins |
Time spent on site before checkout |
converted |
Binary purchase outcome (1 = purchased, 0 = abandoned) |
Embedded Behavioral Patterns:
- Returning customers convert more often than new customers
- Mobile users convert less often than desktop users
- The variant receives a positive treatment effect in the data-generating process
- Missingness is introduced in
time_spent_minsto reflect real-world data quality issues
This design allows the project to test whether observed uplift survives after controlling for user-level characteristics.
| Test | Implementation | Purpose |
|---|---|---|
| Two-proportion z-test | scipy.stats.norm |
Determine whether the conversion rate difference is statistically significant |
| Logistic regression | statsmodels.api.Logit |
Estimate treatment effect after controlling for device type, customer type, and cart value |
| Logistic regression with interactions | statsmodels.api.Logit |
Formally test whether treatment effects differ across user segments (HTE) |
| Power analysis | statsmodels.stats.power.zt_ind_solve_power |
Validate that sample size is sufficient to detect a meaningful effect |
| Bayesian A/B test | Beta-Binomial conjugate model via numpy.random.beta |
Express uncertainty as decision-friendly probabilities |
| A/A validation test | Two-proportion z-test on Control splits | Verify randomization and instrumentation produce no false positives |
| Sample Ratio Mismatch (SRM) | scipy.stats.chisquare |
Confirm observed traffic split matches intended 50/50 assignment |
-
Z-test is the standard frequentist approach for comparing two proportions. It provides a p-value and confidence interval that are directly interpretable for launch decisions.
-
Logistic regression isolates the treatment effect from confounding variables. Even in a randomised experiment, verifying that the effect survives covariate adjustment strengthens the conclusion.
-
Power analysis ensures the experiment was designed to detect a practically meaningful effect before data collection began. This is a critical step that most portfolio projects omit.
-
Bayesian analysis answers the questions product teams actually care about: "What is the probability the variant wins?" and "What is the downside risk?"
| Metric | Value |
|---|---|
| Control conversion rate | 19.68% |
| Variant conversion rate | 24.11% |
| Absolute uplift | 4.43 percentage points |
| Relative uplift | 22.49% |
| Z-statistic | 5.35 |
| P-value (one-sided) | 4.32e-08 |
| 95% CI for uplift | [2.81 pp, 6.05 pp] |
| Variable | Odds Ratio | Interpretation |
|---|---|---|
| Variant (treatment) | 1.33 | Users in the variant are 33% more likely to convert, holding all else constant |
| Returning customer | 1.67 | Returning customers convert at significantly higher rates |
| Mobile device | 0.77 | Mobile users convert less frequently than desktop users |
| Decision Metric | Value |
|---|---|
| P(Treatment beats Control) | >99.99% |
| P(Uplift > 1 pp threshold) | >99% |
| P(Annual Revenue Gain > $1.0M) | 98.93% |
| Estimated posterior probability that variant underperforms control | 0.00% (Near zero) |
| Expected conversion uplift | 4.43 pp |
| 90% credible interval for conversion uplift | [3.06 pp, 5.78 pp] |
| Expected annual revenue gain | $1.76M |
| 90% credible interval for annual revenue gain | [$1.22M, $2.29M] |
To understand if the treatment effect was consistent across cohorts, we performed segment-specific two-proportion Z-tests:
| Cohort Group | Cohort Segment | Control CR | Variant CR | Absolute Uplift | Relative Lift | P-Value (one-sided) | Statistically Significant? |
|---|---|---|---|---|---|---|---|
| Device Type | Mobile | 17.72% | 22.09% | +4.37 pp | +24.67% | 0.000012 | Yes (p < 0.05) |
| Device Type | Desktop | 21.73% | 26.85% | +5.12 pp | +23.58% | 0.000499 | Yes (p < 0.05) |
| Device Type | Tablet | 22.54% | 28.49% | +5.94 pp | +26.36% | 0.014667 | Yes (p < 0.05) |
| Customer Type | New Customers | 16.02% | 17.53% | +1.51 pp | +9.40% | 0.101572 | No (p = 0.101) |
| Customer Type | Returning Customers | 21.66% | 28.67% | +7.01 pp | +32.38% | 0.000000 | Yes (p < 0.001) |
- Mobile and Desktop Universality: The One-Click checkout is highly effective across both Mobile (+4.37 pp uplift) and Desktop (+5.12 pp uplift) devices, showing robust performance. Because Mobile traffic represents 60% of all sessions, this mobile lift is the core volume driver for overall business growth.
-
The Loyalty HTE: Slicing by customer history reveals a classic Heterogeneous Treatment Effect. The variant is extremely effective for Returning Customers, boosting their conversion rate by +7.01 pp (a +32.38% relative lift). However, for New Customers, the conversion uplift is only +1.51 pp and is not statistically significant (
$p = 0.101 > 0.05$ ). - Business Rationale: Returning customers already have saved shipping/billing details and high brand trust, so "One-Click Checkout" completely eliminates purchase friction. New customers do not have pre-filled details, so they must still enter shipping/billing addresses manually, meaning the One-Click interface cannot eliminate their primary entry friction, and their baseline trust remains lower. Future product iterations should target guest-checkout autofills and trust signals to help new customers convert.
While separate segment-level z-tests are informative, they do not formally test whether the treatment effect differs across segments. To do this rigorously, we fit a logistic regression with an interaction term:
| Term | Odds Ratio | 95% CI | P-Value | Interpretation |
|---|---|---|---|---|
| Treatment (Variant) | ~1.10 | [0.87, 1.39] | ~0.42 | Baseline treatment effect for New Customers (reference group) — not significant alone |
| Returning Customer | ~1.44 | [1.18, 1.76] | ~0.0003 | Returning customers convert more regardless of treatment |
| Treatment × Returning | ~1.35 | [1.04, 1.76] | ~0.025 | The treatment effect is significantly stronger for Returning Customers |
The significant interaction term (
Deploy One-Click Checkout universally, but prioritize future optimization efforts for New Customers where uplift remains weak and not statistically significant. Specifically:
- Immediate: Launch the variant for all users — the overall uplift is strong and significant.
- Next Sprint: Investigate guest-checkout autofill (e.g., Shop Pay, Google Pay, Apple Pay) to reduce first-time data-entry friction for new customers.
- Next Quarter: Run a dedicated experiment targeting new-customer trust signals (security badges, guest checkout, social proof) to close the gap.
The 95% confidence interval for the conversion uplift is [2.81 pp, 6.05 pp]. The entire interval lies above zero, meaning:
- The improvement is not just statistically significant; it is directionally reliable.
- Even in the worst-case scenario (lower bound of the CI), the variant still delivers a 2.81 percentage point improvement.
- The Bayesian 90% credible interval [3.06 pp, 5.78 pp] corroborates this finding.
This framing is valuable for executives because it translates statistical uncertainty into a range of plausible business outcomes rather than a single point estimate.
A raw conversion difference can be useful, but it is not always sufficient. The key question is whether the uplift remains after controlling for characteristics that independently affect conversion.
| Metric | Naive Result | Adjusted Result |
|---|---|---|
| Variant uplift | +4.43 pp | +4.47 pp |
| Statistical significance | Strong | Strong |
| Interpretation | Variant appears better | Effect remains after controlling for covariates |
In this experiment, the adjusted effect remains close to the naive estimate because randomisation was clean. This is the expected, reassuring outcome for a well-designed experiment. The comparison is included to demonstrate the methodology and verify that no hidden confounding is inflating the result.
| Assumption | Value |
|---|---|
| Annual visitors | 1,200,000 |
| Median AOV | $33.09 |
| Control annual revenue | ~$7.81M |
| Variant annual revenue | ~$9.57M |
| Incremental revenue | $1.76M per year |
A 4.43 percentage point conversion improvement, applied to 1.2M annual visitors at a $33.09 median order value, projects to approximately $1.76M in additional annual revenue.
Note: Traffic and AOV assumptions are representative of a mid-sized e-commerce platform (comparable to Shopify merchant benchmarks) and are used solely for scenario analysis. In production, these figures should be replaced with actual platform analytics.
This moves the result from "Variant B wins" to "Variant B wins and is worth approximately $1.76M annually." This is what stakeholders care about.
| Result | Action | Rationale |
|---|---|---|
| Significant positive uplift | Deploy variant | Evidence supports improved conversion |
| No significant difference | Keep control | Insufficient evidence to justify change |
| Significant negative uplift | Rollback to control | Variant harms conversion |
This experiment result: Significant positive uplift. Recommendation is to deploy.
Most portfolio projects stop at reporting a p-value. A rigorous analysis also considers the error risks inherent in the decision.
- Definition: Concluding the variant is better when it is not (rejecting H0 when H0 is true).
- Control: The significance level alpha = 0.05 limits the false positive rate to 5%.
- In this experiment: The observed p-value (4.32e-08) is far below alpha, providing extremely strong evidence against the null hypothesis.
- Definition: Failing to detect a real improvement (failing to reject H0 when H1 is true).
- Control: The experiment was designed with 80% power (beta = 0.20), meaning a 20% chance of missing a real 15% relative lift.
- In this experiment: The observed effect (22.49% relative lift) far exceeds the minimum detectable effect (15%), so the risk of a Type II error is negligible for this effect size.
| Risk | Probability / Control | Consequence | Mitigation |
|---|---|---|---|
| False positive (alpha) | Controlled at 5% (Observed p-value << alpha provides strong evidence against H0) | Launching a feature that does not actually improve conversion | Pre-registered significance level; post-launch monitoring |
| False negative (beta) | Negligible at observed effect size | Missing a genuine improvement | Adequate sample size via power analysis |
This analysis demonstrates that the experimental design properly controls for both error types, and the observed results fall well within the regime where both risks are minimal.
Frequentist testing tells us whether the uplift is statistically significant. Bayesian analysis answers the questions product teams usually care about most:
- What is the probability that the variant beats the control? Greater than 99.99%.
- What is the probability that the uplift exceeds a business threshold? High (>99% for a 1 pp threshold).
- What is the probability that the annual revenue gain exceeds $1.0M? 98.93%.
- What is the downside risk if the feature is launched? The estimated posterior probability that the variant underperforms the control is 0.00% (near zero).
This makes the result easier to communicate in decision language rather than only p-values.
Visualizations are essential to communicate complex statistical findings to business stakeholders. Below are the key visual assets from the analysis with direct business interpretations.
Interpretation: The error bars show the 95% confidence intervals for baseline conversion rates and the resulting treatment uplift; because the entire uplift interval lies strictly above zero, the conversion gain is statistically valid and directionally reliable.
Interpretation: The odds ratio plot displays the relative impact of each user trait on conversion; being in the Variant group increases a user's likelihood of conversion by ~33%, while returning customers convert at a significantly higher rate, and mobile users convert less frequently.
Interpretation: The annual revenue impact simulation maps the absolute conversion lift against our 1.2 million visitor volume and $33.09 median AOV, projecting an expected incremental revenue of $1.76 million per year (with a 95% range from $1.11M to $2.40M).
The dataset is entirely simulated using np.random.seed(42). No real customer transactions were used. The methodology is fully transferable to production data, but the specific numerical results should be interpreted as a demonstration of the analytical framework.
-
Short experiment duration: A 2-week observation window may capture novelty effects that decay over time. Post-launch monitoring over 6+ weeks is recommended to establish the true long-term conversion plateau.
-
Limited traffic volume: The simulation uses 10,000 users. Real-world experiments at larger scale may reveal effects not visible in this sample.
-
Seasonality effects not considered: Conversion rates fluctuate by day-of-week, time-of-day, and season. The simulation assumes a static data-generating process.
-
User behavior may change post-deployment: The Hawthorne effect and novelty bias can inflate initial results. Long-term holdout groups would help quantify this.
-
Cannibalization risk: Conversion was measured, but average order value should be verified separately to ensure users are not checking out prematurely (before adding secondary items).
-
No network or social effects: Users in production may influence each other through word-of-mouth or shared devices, violating the independence assumption (SUTVA).
-
Revenue projections depend on assumptions: The $1.76M figure assumes 1.2M annual visitors and a $33.09 median AOV, representative of a mid-sized e-commerce platform. These should be replaced with actual platform analytics before making investment decisions.
Because this project uses simulated data with an embedded positive treatment effect, every statistical test converges on the same conclusion: the variant wins. In practice, A/B test outcomes are rarely this clean. To demonstrate analytical maturity, it is important to acknowledge the realistic scenarios where the outcome would differ:
The variant might increase conversion rate while decreasing average order value. If the One-Click flow encourages impulsive, low-value purchases (users checking out before adding secondary items), the net Revenue Per Visitor could decline despite more transactions. This is why AOV and RPV are tracked as guardrail metrics.
The variant might help desktop users but hurt mobile users — for example, if the One-Click UI introduces a mobile rendering bug or eliminates a trust-building step that mobile users rely on. The segmented HTE analysis is specifically designed to catch this pattern.
With a smaller sample size or a smaller true effect, the experiment might fail to reach statistical significance entirely. The power analysis (Section 5) was designed to prevent this, but real-world constraints (limited traffic, shorter test windows, smaller true effects) frequently produce inconclusive results where the confidence interval spans zero.
The initial uplift might be driven entirely by novelty — returning users engage more because the flow is new, not because it is better. A longer observation window or post-launch holdout would reveal whether the effect decays to zero over 4–6 weeks.
In a production experimentation program, roughly 70–90% of A/B tests produce null or inconclusive results (source: Microsoft, Google, Booking.com published research). A project that only shows a clean win risks appearing engineered. The methodology in this project — power analysis, SRM checks, covariate adjustment, Bayesian risk quantification, and post-launch monitoring — is designed to produce trustworthy conclusions regardless of whether the outcome is positive, negative, or null.
-
Multi-armed bandits: Explore Thompson Sampling or UCB algorithms to dynamically allocate traffic during the experiment, reducing opportunity cost.
-
Sequential testing: Implement group sequential designs or always-valid p-values to allow early stopping without inflating Type I error.
-
Long-term holdout groups: Maintain a small percentage of users on the control experience post-launch to measure long-term treatment effects and novelty decay.
-
Heterogeneous treatment effects: Use causal forests or interaction terms to identify user segments where the treatment effect is strongest or weakest.
-
Metric sensitivity analysis: Evaluate alternative primary metrics (revenue per visitor, time-to-purchase) to ensure the conversion lift does not come at the expense of other business objectives.
A statistically significant experiment result is only the beginning. Real product impact must be validated through disciplined post-launch monitoring. Below is the phased plan we would follow after deploying the One-Click Checkout to 100% of traffic:
| Phase | Timeline | Metrics to Monitor | Goal |
|---|---|---|---|
| Phase 1: Stability Check | Week 1–2 | Conversion rate, checkout error rate, payment failure rate, page load time | Confirm the variant performs as expected in production with no technical regressions. |
| Phase 2: Novelty Decay | Week 3–6 | Daily conversion rate trend, returning-user conversion delta | Detect whether the initial uplift decays as the "new feature" novelty wears off. Establish the true long-term conversion plateau. |
| Phase 3: Guardrail Validation | Month 2 | Average order value, refund/chargeback rate, revenue per visitor | Verify that improved conversion has not come at the expense of downstream business metrics. |
| Phase 4: Retention & LTV | Month 3–6 | 30-day repeat purchase rate, customer lifetime value (LTV), cohort retention curves | Evaluate whether the faster checkout experience improves (or harms) long-term customer loyalty and repeat purchase behavior. |
| Phase 5: Segment Deep-Dive | Ongoing | New-customer conversion, mobile-specific funnel metrics | Track whether future product iterations (guest autofill, trust signals) successfully close the gap for new customers identified in the HTE analysis. |
Holdout Group: Maintain a 5% random holdout on the control experience for at least 90 days post-launch. This provides a clean causal baseline to measure long-term effects and detect novelty decay without confounding.
conversion-optimization-ab-testing/
├── README.md
├── Makefile
├── requirements.txt
├── analysis/
│ └── Conversion_Optimization_Analysis.ipynb
├── src/
│ ├── bayesian_ab.py
│ ├── frequentist_ab.py
│ ├── hte_analysis.py
│ └── power_analysis.py
├── scripts/
│ ├── create_notebook.py
│ └── enhance_notebook.py
├── assets/
│ ├── conversion_rate.png
│ ├── odds_ratio_plot.png
│ └── revenue_impact.png
└── .gitignore
-
Python
-
NumPy, pandas for data preparation
-
Matplotlib, seaborn for visualization
-
SciPy for statistical testing
-
statsmodels for logistic regression and power analysis
-
scikit-learn for preprocessing
Sanman Kadam MSc Statistics | Aspiring Data Analyst / Data Scientist
GitHub: https://github.com/the-irritater
LinkedIn: https://www.linkedin.com/in/sanman-kadam-7a4990374/
Created as a portfolio project demonstrating applied A/B testing, causal reasoning, and business decision-making.