{"title":"Evidence, not a verdict.","description":"A p-value cannot tell you whether an effect matters. Explore sample size, uncertainty and the hidden cost of testing everything.","theme":"rose","kind":"Statistical testing","cases":[{"id":"effect-0","name":"No effect · A = B","title":"Evidence depends on the sample.","unit":"Fraction of simulated experiments · 0–1","parameter":"Visitors per arm","x":[100,200,300,500,750,1000,1500,2000,3000,5000],"series":[{"label":"Two-sided detection","values":[0.0516667,0.049,0.056,0.0463333,0.0466667,0.0543333,0.049,0.0473333,0.048,0.0486667],"low":[0.0443035,0.0418356,0.048326,0.039374,0.0396813,0.0467772,0.0418356,0.0402963,0.0409117,0.0415275],"high":[0.0601766,0.0573179,0.0648096,0.054453,0.0548115,0.0630294,0.0573179,0.0555281,0.0562444,0.0569602]},{"label":"Positive & significant","values":[0.025,0.0253333,0.029,0.0243333,0.0223333,0.0213333,0.0246667,0.0263333,0.024,0.026],"low":[0.0199913,0.0202883,0.0235713,0.019398,0.0176248,0.0167421,0.0196946,0.0211809,0.0191018,0.0208831],"high":[0.0312236,0.0315924,0.0356334,0.0304852,0.0282636,0.0271488,0.0308545,0.0326972,0.0301157,0.0323292]},{"label":"95% interval coverage","values":[0.9536667,0.951,0.9473333,0.9536667,0.9543333,0.9456667,0.951,0.9533333,0.952,0.9516667],"low":[0.945547,0.9426821,0.938753,0.945547,0.9462642,0.9369706,0.9426821,0.9451885,0.9437556,0.9433977],"high":[0.960626,0.9581644,0.9547696,0.960626,0.9612404,0.9532228,0.9581644,0.9603187,0.9590883,0.9587804]}],"snapshots":[{"rows":[["Visitors per arm",100],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.004 percentage points"],["Mean interval width","17.210 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",200],["True absolute effect","0.0 percentage points"],["Mean observed effect","0.057 percentage points"],["Mean interval width","11.978 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",300],["True absolute effect","0.0 percentage points"],["Mean observed effect","0.037 percentage points"],["Mean interval width","9.712 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",500],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.014 percentage points"],["Mean interval width","7.493 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",750],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.021 percentage points"],["Mean interval width","6.105 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1000],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.029 percentage points"],["Mean interval width","5.284 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1500],["True absolute effect","0.0 percentage points"],["Mean observed effect","0.011 percentage points"],["Mean interval width","4.307 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",2000],["True absolute effect","0.0 percentage points"],["Mean observed effect","0.011 percentage points"],["Mean interval width","3.726 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",3000],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.012 percentage points"],["Mean interval width","3.042 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",5000],["True absolute effect","0.0 percentage points"],["Mean observed effect","-0.006 percentage points"],["Mean interval width","2.354 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."}],"columns":["Measurement","Value"],"context":"Control converts at 10%; treatment at 10%. Each sample size contains 3,000 independently simulated experiments.","readout":"Two-sided pooled z-test, α=0.05. Coverage uses a Newcombe interval for the risk difference. Shading is a Wilson 95% interval for Monte Carlo frequencies, not an interval for the conversion uplift.","note":""},{"id":"effect-2","name":"Small effect · +2 percentage points","title":"Evidence depends on the sample.","unit":"Fraction of simulated experiments · 0–1","parameter":"Visitors per arm","x":[100,200,300,500,750,1000,1500,2000,3000,5000],"series":[{"label":"Two-sided detection","values":[0.079,0.1016667,0.125,0.1723333,0.2276667,0.3046667,0.4093333,0.5186667,0.7016667,0.9123333],"low":[0.0698773,0.0913568,0.113643,0.15924,0.2130154,0.288455,0.3918648,0.5007747,0.6850451,0.9016787],"high":[0.0891995,0.1129954,0.1373161,0.1862647,0.2430145,0.3213779,0.4270337,0.5365108,0.7177724,0.9219333]},{"label":"Positive & significant","values":[0.0683333,0.097,0.1206667,0.1706667,0.2266667,0.304,0.4093333,0.5186667,0.7016667,0.9123333],"low":[0.0598454,0.0869191,0.1094929,0.1576273,0.2120399,0.2877993,0.3918648,0.5007747,0.6850451,0.9016787],"high":[0.0779253,0.1081117,0.1328106,0.1845483,0.2419925,0.320702,0.4270337,0.5365108,0.7177724,0.9219333]},{"label":"95% interval coverage","values":[0.9486667,0.9533333,0.9506667,0.9536667,0.9533333,0.9493333,0.9503333,0.9493333,0.9466667,0.957],"low":[0.9401804,0.9451885,0.9423244,0.945547,0.9451885,0.9408947,0.9419668,0.9408947,0.9380398,0.9491377],"high":[0.9560053,0.9603187,0.9578563,0.960626,0.9603187,0.9566227,0.957548,0.9566227,0.9541511,0.9636934]}],"snapshots":[{"rows":[["Visitors per arm",100],["True absolute effect","2.0 percentage points"],["Mean observed effect","2.051 percentage points"],["Mean interval width","17.813 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",200],["True absolute effect","2.0 percentage points"],["Mean observed effect","2.000 percentage points"],["Mean interval width","12.435 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",300],["True absolute effect","2.0 percentage points"],["Mean observed effect","2.071 percentage points"],["Mean interval width","10.114 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",500],["True absolute effect","2.0 percentage points"],["Mean observed effect","1.980 percentage points"],["Mean interval width","7.798 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",750],["True absolute effect","2.0 percentage points"],["Mean observed effect","1.989 percentage points"],["Mean interval width","6.358 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1000],["True absolute effect","2.0 percentage points"],["Mean observed effect","2.019 percentage points"],["Mean interval width","5.498 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1500],["True absolute effect","2.0 percentage points"],["Mean observed effect","1.965 percentage points"],["Mean interval width","4.486 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",2000],["True absolute effect","2.0 percentage points"],["Mean observed effect","1.987 percentage points"],["Mean interval width","3.883 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",3000],["True absolute effect","2.0 percentage points"],["Mean observed effect","1.994 percentage points"],["Mean interval width","3.168 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",5000],["True absolute effect","2.0 percentage points"],["Mean observed effect","2.034 percentage points"],["Mean interval width","2.453 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."}],"columns":["Measurement","Value"],"context":"Control converts at 10%; treatment at 12%. Each sample size contains 3,000 independently simulated experiments.","readout":"Two-sided pooled z-test, α=0.05. Coverage uses a Newcombe interval for the risk difference. Shading is a Wilson 95% interval for Monte Carlo frequencies, not an interval for the conversion uplift.","note":""},{"id":"effect-5","name":"Larger effect · +5 percentage points","title":"Evidence depends on the sample.","unit":"Fraction of simulated experiments · 0–1","parameter":"Visitors per arm","x":[100,200,300,500,750,1000,1500,2000,3000,5000],"series":[{"label":"Two-sided detection","values":[0.1913333,0.3086667,0.448,0.675,0.8333333,0.9326667,0.99,0.9986667,1.0,1.0],"low":[0.1776559,0.29239,0.4302828,0.6580252,0.8195729,0.9231346,0.9857604,0.9965765,0.9987212,0.9987212],"high":[0.2058002,0.3254327,0.4658502,0.6915272,0.8462412,0.9410921,0.9929863,0.9994814,1.0,1.0]},{"label":"Positive & significant","values":[0.191,0.3083333,0.4476667,0.675,0.8333333,0.9326667,0.99,0.9986667,1.0,1.0],"low":[0.1773324,0.292062,0.4299512,0.6580252,0.8195729,0.9231346,0.9857604,0.9965765,0.9987212,0.9987212],"high":[0.205458,0.3250949,0.465516,0.6915272,0.8462412,0.9410921,0.9929863,0.9994814,1.0,1.0]},{"label":"95% interval coverage","values":[0.95,0.9526667,0.957,0.9506667,0.9453333,0.9556667,0.9526667,0.947,0.954,0.948],"low":[0.9416094,0.9444719,0.9491377,0.9423244,0.9366144,0.9477001,0.9444719,0.9383963,0.9459055,0.9394665],"high":[0.9572397,0.9597037,0.9636934,0.9578563,0.9529132,0.9624678,0.9597037,0.9544604,0.9609333,0.9553876]}],"snapshots":[{"rows":[["Visitors per arm",100],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.968 percentage points"],["Mean interval width","18.683 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",200],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.878 percentage points"],["Mean interval width","13.050 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",300],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.918 percentage points"],["Mean interval width","10.635 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",500],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.979 percentage points"],["Mean interval width","8.213 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",750],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.976 percentage points"],["Mean interval width","6.696 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1000],["True absolute effect","5.0 percentage points"],["Mean observed effect","5.008 percentage points"],["Mean interval width","5.795 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",1500],["True absolute effect","5.0 percentage points"],["Mean observed effect","5.012 percentage points"],["Mean interval width","4.727 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",2000],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.998 percentage points"],["Mean interval width","4.093 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",3000],["True absolute effect","5.0 percentage points"],["Mean observed effect","5.004 percentage points"],["Mean interval width","3.340 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."},{"rows":[["Visitors per arm",5000],["True absolute effect","5.0 percentage points"],["Mean observed effect","4.980 percentage points"],["Mean interval width","2.587 percentage points"],["Repeated experiments",3000]],"note":"With no true effect, detection is a false-positive rate—not power. With a positive effect, detection estimates two-sided power. Not rejecting the null does not demonstrate equivalence. These are fixed-horizon experiments; do not stop when p first drops below 0.05."}],"columns":["Measurement","Value"],"context":"Control converts at 10%; treatment at 15%. Each sample size contains 3,000 independently simulated experiments.","readout":"Two-sided pooled z-test, α=0.05. Coverage uses a Newcombe interval for the risk difference. Shading is a Wilson 95% interval for Monte Carlo frequencies, not an interval for the conversion uplift.","note":""},{"id":"multiplicity","name":"Many tests · all null hypotheses true","title":"More chances to fool yourself.","unit":"Probability of ≥1 false rejection","parameter":"Independent tests","x":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50],"series":[{"label":"Unadjusted","values":[0.0513333,0.103,0.1523333,0.1956667,0.2316667,0.2673333,0.299,0.33,0.3623333,0.3953333,0.422,0.4543333,0.4806667,0.505,0.5316667,0.5546667,0.5786667,0.5986667,0.6176667,0.6383333,0.6553333,0.672,0.688,0.7,0.7136667,0.7256667,0.7366667,0.75,0.7633333,0.7776667,0.788,0.802,0.8116667,0.8226667,0.835,0.8453333,0.852,0.8593333,0.865,0.871,0.8756667,0.883,0.889,0.894,0.9003333,0.9056667,0.909,0.9143333,0.9196667,0.9236667],"low":[0.0439947,0.092626,0.1399198,0.1818637,0.2169185,0.2518014,0.2828829,0.3134007,0.3453191,0.3779823,0.4044379,0.4365859,0.4628242,0.487114,0.5137815,0.5368233,0.560908,0.5810111,0.6001372,0.620973,0.6381378,0.6549894,0.6711894,0.6833545,0.6972255,0.7094197,0.7206105,0.734192,0.7477931,0.7624374,0.7730106,0.7873581,0.7972806,0.8085889,0.8212909,0.8319535,0.8388431,0.846432,0.8523039,0.8585291,0.8633767,0.8710055,0.8772578,0.882476,0.8890968,0.8946824,0.8981785,0.903781,0.909395,0.9136138],"high":[0.0598196,0.1143894,0.1656361,0.2102481,0.2471012,0.2834603,0.3156312,0.3470341,0.3796997,0.412952,0.4397616,0.4721976,0.4985585,0.5228732,0.5494708,0.5723702,0.5962242,0.6160698,0.6348952,0.6553399,0.6721316,0.6885707,0.7043297,0.716134,0.7295614,0.7413364,0.7521175,0.7651686,0.7782,0.7921857,0.8022527,0.8158695,0.8252555,0.8359192,0.8478523,0.8578299,0.8642566,0.8713156,0.8767625,0.882522,0.8869958,0.8940149,0.8997472,0.9045162,0.910546,0.9156134,0.9187754,0.923826,0.9288649,0.9326359]},{"label":"Holm","values":[0.0513333,0.048,0.0543333,0.0523333,0.0513333,0.048,0.0496667,0.0486667,0.047,0.0456667,0.045,0.047,0.0463333,0.0443333,0.042,0.0413333,0.042,0.042,0.0436667,0.0426667,0.0433333,0.042,0.0426667,0.044,0.0453333,0.0456667,0.0446667,0.044,0.045,0.044,0.0436667,0.0456667,0.045,0.046,0.046,0.0466667,0.047,0.0473333,0.0473333,0.048,0.0476667,0.0483333,0.0483333,0.0476667,0.0476667,0.0476667,0.0476667,0.047,0.048,0.048],"low":[0.0439947,0.0409117,0.0467772,0.0449214,0.0439947,0.0409117,0.042452,0.0415275,0.0399888,0.0387596,0.0381457,0.0399888,0.039374,0.0375322,0.0353886,0.0347772,0.0353886,0.0353886,0.0369191,0.0360004,0.0366128,0.0353886,0.0360004,0.0372256,0.0384526,0.0387596,0.0378389,0.0372256,0.0381457,0.0372256,0.0369191,0.0387596,0.0381457,0.0390667,0.0390667,0.0396813,0.0399888,0.0402963,0.0402963,0.0409117,0.040604,0.0412196,0.0412196,0.040604,0.040604,0.040604,0.040604,0.0399888,0.0409117,0.0409117],"high":[0.0598196,0.0562444,0.0630294,0.0608903,0.0598196,0.0562444,0.0580332,0.0569602,0.0551699,0.0537358,0.0530181,0.0551699,0.054453,0.0522999,0.0497829,0.0490626,0.0497829,0.0497829,0.0515814,0.0505026,0.0512219,0.0497829,0.0505026,0.0519407,0.053377,0.0537358,0.0526591,0.0519407,0.0530181,0.0519407,0.0515814,0.0537358,0.0530181,0.0540945,0.0540945,0.0548115,0.0551699,0.0555281,0.0555281,0.0562444,0.0558863,0.0566023,0.0566023,0.0558863,0.0558863,0.0558863,0.0558863,0.0551699,0.0562444,0.0562444]},{"label":"Benjamini–Hochberg","values":[0.0513333,0.0493333,0.055,0.0533333,0.0516667,0.0486667,0.0506667,0.05,0.0486667,0.0463333,0.0463333,0.048,0.0473333,0.0453333,0.044,0.043,0.043,0.0423333,0.0443333,0.043,0.0436667,0.0433333,0.044,0.045,0.0463333,0.0463333,0.0453333,0.0443333,0.0453333,0.0443333,0.0446667,0.0466667,0.046,0.0466667,0.0466667,0.0476667,0.048,0.0486667,0.0483333,0.049,0.0486667,0.0496667,0.0496667,0.0486667,0.0486667,0.0486667,0.049,0.0483333,0.049,0.049],"low":[0.0439947,0.0421437,0.0473964,0.0458489,0.0443035,0.0415275,0.0433773,0.0427603,0.0415275,0.039374,0.039374,0.0409117,0.0402963,0.0384526,0.0372256,0.0363066,0.0363066,0.0356944,0.0375322,0.0363066,0.0369191,0.0366128,0.0372256,0.0381457,0.039374,0.039374,0.0384526,0.0375322,0.0384526,0.0375322,0.0378389,0.0396813,0.0390667,0.0396813,0.0396813,0.040604,0.0409117,0.0415275,0.0412196,0.0418356,0.0415275,0.042452,0.042452,0.0415275,0.0415275,0.0415275,0.0418356,0.0412196,0.0418356,0.0418356],"high":[0.0598196,0.0576756,0.0637417,0.0619602,0.0601766,0.0569602,0.0591053,0.0583906,0.0569602,0.054453,0.054453,0.0562444,0.0555281,0.053377,0.0519407,0.0508623,0.0508623,0.0501428,0.0522999,0.0508623,0.0515814,0.0512219,0.0519407,0.0530181,0.054453,0.054453,0.053377,0.0522999,0.053377,0.0522999,0.0526591,0.0548115,0.0540945,0.0548115,0.0548115,0.0558863,0.0562444,0.0569602,0.0566023,0.0573179,0.0569602,0.0580332,0.0580332,0.0569602,0.0569602,0.0569602,0.0573179,0.0566023,0.0573179,0.0573179]}],"context":"3,000 families of independent, uniform null p-values. Each family grows from 1 to 50 tests.","readout":"Here every null is true, so any rejection is false. Holm's first threshold is α/m. BH is evaluated with its step-up rule. Shading quantifies simulation error only.","note":"Holm controls family-wise error. BH controls false discovery rate under independence; only in this all-null setting does FDR equal the chance of any rejection. This does not make BH a general FWER procedure."}],"config":{"seed":20260927,"multiplicitySeed":77,"repetitions":3000,"alpha":0.05,"scipy":"1.18.1","sizes":[100,200,300,500,750,1000,1500,2000,3000,5000]},"method":["Binary outcomes are independent Bernoulli draws with equal sample sizes in each arm. Each experiment is analyzed once at its planned horizon. The pooled two-proportion z-test is two-sided at α=0.05.","The risk-difference interval combines separate Wilson intervals using Newcombe's uncorrected method. Test and interval use different approximations, so they are not guaranteed to give identical rejection decisions.","Each point summarizes 3,000 simulated experiments. Wilson intervals on frequency curves quantify finite Monte Carlo uncertainty. Multiple-testing examples use independent uniform null p-values and a separate seed."],"limits":["These are synthetic teaching examples, not an analysis of customer experiments or clinical advice. Unequal allocation, repeated users, clustering and optional stopping require different designs.","A significant result need not be valuable; a non-significant result need not be zero. Specify the minimum worthwhile effect before collecting data. The sparse-data case may require exact tests."],"references":[["SciPy · Statistical functions","https://docs.scipy.org/doc/scipy/reference/stats.html"],["Newcombe · Two independent proportions (1998)","https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8%3C873::AID-SIM779%3E3.0.CO;2-I"]],"provenance":{"python":"3.13.0","numpy":"2.4.6","sourceSha256":{"showcase_benchmark.py":"5b2cb15798cb5fbfc0696ecbf2fe60129a52ca0b271ac467fa7c4108b79bb4b0"}}}
