chile-fintech-experimentation-lab

A/B testing done the way a real experimentation platform has to do it: power analysis before launch, Sample Ratio Mismatch (SRM) detection, CUPED variance reduction, multiple-testing correction across guardrail metrics, a novelty-effect check — and, as the centerpiece, an empirical calibration harness that runs each stopping rule thousands of times under a *known* ground truth to check whether it actually controls the error rate it claims to.

This page shows the project's results. The methodology, the design decisions and the limitations are documented in the repository's README.

chile-fintech-experimentation-lab

False-positive rate against the number of daily looks
At a single look the test delivers what it promises (3.8%, consistent with the nominal 5%). The damage is done by the decision to keep looking, and it accumulates with each additional day.

1. Power analysis

Required sample size against minimum detectable effect
Verified against a known textbook case in tests/test_power.py: baseline 20%, MDE +5pp absolute, alpha=0.05, power=0.80 reproduces 1,091 per arm, matching Evan Miller's public A/B testing sample-size calculator for the identical input, hand-checked from the pooled-variance formula. That validation point is the blue marker on the chart.

3. CUPED variance reduction

CUPED variance reduction and its effect on the confidence interval
41.9% variance reduction using a pre-period revenue covariate. The pipeline specifies a correlation of 0.65; the realized correlation in the sample is ρ = 0.6474, so CUPED's theoretical ceiling here is ρ² = 41.92% — and the measured reduction is 41.9168%, sitting exactly on the ceiling rather than short of it. (The reduction CUPED achieves *is* the empirical ρ², which is why the two agree to four decimals. The only gap worth naming is between the realized ρ and the 0.65 specified, which is ordinary sampling error in the correlation itself, not slack in the method.)

4. Multiple-testing correction (BH-FDR) across 7 guardrail metrics

Guardrail p-values against the BH critical line
Read at a raw alpha=0.05, 4 of 7 metrics look significant. Under BH-FDR control, only 1 survives. Three of those four (latency, unsubscribe, support tickets) are exactly the guardrail metrics a launch decision should be most cautious about overreacting to — this is what the correction is for.

5. Novelty-effect check

Injected exponential decay against the fitted linear interaction
Interaction coefficient -0.00433, p < 0.0001 — correctly detects an injected decaying effect.

6. Calibration: does each stopping rule control what it claims to?

Empirical false-positive rate of each stopping rule against its nominal 5%
One caveat the figure makes visible and the table does not: the two Wilson intervals overlap ([20.3%, 28.7%] and [16.8%, 24.7%]). At 400 simulations the Bayesian rule is *not* distinguishable from the naive z-test — both are distinguishable from the 5% they claim, which is the finding, but "Bayesian is better than naive" is not something this evidence supports. Separating them would need more simulations.