Test our statistics against your data shape.
Build a synthetic experiment in 60 seconds — your choice of n, effect size, distribution, outliers, and missing-data pattern. Click 'Test it' to run the dataset through a faithful, illustrative approximation of Winnow's statistical methods. We surface every method that fired and how we reasoned. No login. No data ever uploaded.
Two variants. One question: does treatment beat control?
Configure baseline mean, effect size, distribution (normal / lognormal / heavy-tail), outliers, and missing-data pattern (MCAR / MAR / MNAR). Watch a histogram update live as you tweak. Click 'Test it' to run mSPRT with auto-cure CUPED.
Two prompts. One question: which scores higher on a judge?
Configure baseline pass rate, treatment uplift, refusal asymmetry, and missing pattern. MNAR is the headliner — when the worst answers silently fail to log, naive comparisons mislead. Watch our stats refuse to call a winner unless the data actually supports one.
statistical/ engine that runs your production experiments, and the numbers are approximations meant to build intuition, not production results.