It’s a Tuesday morning. The kettle hisses — the new kettle, which boils in half the time and cost him ₹3,100.
Viji bought it six weeks ago on the theory that a faster kettle means a shorter queue, and a shorter queue means nobody walks off to the stall near the bus stand. He has 30 days of customer counts from before it arrived and 30 from after.
Before: mean
87customers a day. After: mean95. SD about13in both stretches.
He runs the two-sample t-test he learned from the last two visitors. The output is one line:
t = 2.383,df = 58,p = 0.020
He knows the rule. 0.020 is less than 0.05, so he rejects the null and concludes the kettle did something.
His cousin, who has developed an unfortunate habit of asking the second question: “And if you had wanted to be strict, brother? If you had used 0.01 instead of 0.05?”
Viji looks at the number. 0.020 is not less than 0.01.
He types back, “Then I wouldn’t reject.”
“So which is it?”
Viji puts the phone down. “It’s the same kettle,” he says, to nobody. “The same 60 days. The same 0.020. How can the answer depend on a number I get to pick?”
The awning rustles.
A woman steps in under the awning — deliberate, unhurried, slate-blue shirt, with a builder’s chalk line reel clipped to her belt. She does not sit. She walks out to the road in front of the stall, kneels, pulls the string taut across the dust, and snaps it. A crisp blue line appears on the ground.
She says, “I am Alpha. I am that line. I am drawn before anything is thrown.”
A man walks in behind her — quick, cheerful, burnt-orange shirt, a steel measuring tape hooked on one finger. He does not look at the line at all. He is looking down the road, at where something has already landed.
He says, “I am the p-value. I measure after. Somebody throws, the stone lands, and I walk out with my tape and report how far from the mark it fell. That is the whole of my job, and I have no opinion at all about whether the distance is good enough.”
Viji looks between them. “But the two of you together decide whether my kettle worked.”
They reply, together, “Only in that order.”
Alpha speaks first
Alpha stands up and dusts off her hands. The blue line sits in the road, plain and permanent.
She says, “Here is what I actually am, and it is not what most people think. I am not a fact about your data. I have never seen your data. I am a policy — the error rate you have decided in advance that you are willing to live with.”
She says, “Set me at 0.05 and you are saying: if the kettle had done nothing at all, I accept a 5% chance of being fooled into saying it did. That is a statement about the procedure, made before the first day was recorded. It stays true whether your p comes out at 0.001 or 0.9.”
Viji says, “Then why can I choose it?”
“Because how much of that mistake you can afford is a question about your stall, not about your statistics. A 5% false-alarm rate on a kettle is nothing — you buy a kettle, and if it turns out to be useless, you own a slightly better kettle. If you were deciding whether to shut the stall and move to another road, you would want me much smaller, because that mistake costs you your living.”
She adds, with some emphasis, “And this is my one rule. I am drawn before the throw. Once the stone is in the air I do not move, and I am not open to negotiation. The moment I am chosen after seeing where it landed, I stop being an error rate and become a decoration.”
The p-value speaks
The p-value unhooks his tape and lets it run out along the road.
He says, “Now me. Everyone says my name and almost nobody says what I measure, so here it is, carefully.”
He says, “I assume the null is true. Not that it is — I assume it, for the sake of measuring. I imagine the kettle did nothing, that your 95 and your 87 differ only because 60 particular days fell the way they did. Then I ask one question: in that imaginary world, how often would I see a gap at least as big as 8 cups?”
“The answer is 0.020. Two times in a hundred.”
Viji says, “So there’s a 2% chance the kettle did nothing?”
“No. And that sentence is the single most common mistake in applied statistics, so let me be exact. I am not the probability that the null is true. I am the probability of data like yours, given that the null is true. Those are different questions, and swapping them is like confusing how often clouds appear when it rains with how often it rains when there are clouds.”
He walks the tape back in.
“Second thing about me, and it is the one that surprises people. I am not a measure of how big the effect is. Take your same 8-cup difference, and your same 13-cup spread, and collect only 10 days on each side instead of 30.”
He writes it in the dust: t = 1.376, p = 0.186.
Viji stares. “The same 8 cups. The same kettle. And now p is 0.186?”
“The same gap, measured with fewer days. I am part evidence and part sample size, and I never tell you which part is which. That is why a tiny effect in a huge study can post a thrilling p, and a large effect in a small study can post a dull one.”
He points at the chalk line. “Her line does not move when n changes. Mine does. We are built out of different material.”
What just happened
Viji pours a glass of tea, sits on his own counter, and writes the two of them out.
| Alpha (α) | p-value | |
|---|---|---|
| Chosen or computed? | chosen, by you | computed, from the data |
| When? | before the data | after the data |
| What is it about? | the procedure | this one sample |
| Depends on n? | no | yes |
| Changes if the data changes? | no | yes |
| Meaning | error rate you accept | P(data this extreme given H₀) |
| The decision | reject if p ≤ α |
He looks at the last row for a long time.
He says, “So the cousin’s question was not really a question about statistics. He asked what I would have concluded at α = 0.01. But I had already chosen 0.05, weeks ago, before the kettle arrived. The honest answer is that 0.01 was never my line.”
Alpha says, “Correct. And if you had chosen 0.01 back then, you would have chosen it for a reason — a more expensive mistake — and you would be living with not rejecting today, and that would also be honest.”
The p-value adds, “What you must not do is look at my 0.020, notice it clears one bar and not the other, and then decide which bar you were aiming at.”
The same chat, in a chart
That picture is the same conversation, drawn. The first panel is the line, snapped before anyone looks. The second is the tape, run out after the stone lands. The inset is the part worth remembering: the stone travelled exactly as far, the line never moved, and the number changed anyway, because fewer days were used to measure it.
One last warning before they leave
The p-value clips his tape shut.
He says, “Three traps, and every one of them is somebody moving her line after the fact.”
“One. Do not choose α after seeing p. It has a name — and if you go looking for a threshold your result happens to clear, your stated error rate is fiction. The 0.05 has to be older than the data.”
“Two. Do not keep collecting until I behave. If you test at 30 days, then add 10 more because I sat at 0.08, then check again, you have taken several shots at the line while only admitting to one. Under repeated peeking, a nominal 5% false-alarm rate can rise past 20%. Fix the sample size in advance, or use a design that is built for interim looks.”
Alpha says, “Three, and this one is mine. 0.05 is a convention, not a law. Ronald Fisher suggested it as a convenient rule of thumb and said plainly that a worker should set his own level for his own purpose. Treating p = 0.049 as a discovery and p = 0.051 as nothing is treating a hand-drawn line as a cliff edge.”
She adds, “So report me, report him, and report the thing neither of us measures: the size of the difference, with a confidence interval around it. 8 more cups a day, somewhere between about 1 and 15, is a sentence your supplier can act on. p < 0.05 is not.”
Viji writes all three down. Draw the line first. Never move it. And say how big the difference was.
The bill
They left the way people leave a tea stall on a working morning. Alpha wound her chalk line back onto its reel and stepped over the blue mark on her way out without looking at it. The p-value coiled his tape, tapped it twice against his palm, and followed her.
Viji kept the kettle. He wrote the decision in the back of his notebook the way they told him to:
α = 0.05, fixed before the kettle arrived. Observed+8cups a day,95%CI roughly1to15.p = 0.020. Rejected. Kettle stays.
Then he added the line he expected to need more often than the rest of it: p is not the chance I am wrong. It is how surprising my days would be if the kettle had done nothing at all.
His cousin, inevitably: “And if I still think 0.01 is the right bar?”
Viji typed back, “Then say so before the next thing I buy, and I’ll test it at 0.01. You don’t get to pick after you’ve seen the number.”
The cousin’s reply took some time. “Fair.”
For the math-curious
The definition. For an observed statistic
T = tand a null hypothesisH₀, the two-sided p-value is $$ p = P\big(|T| \ge |t| ;\big|; H_0\big) $$ It is a tail probability under the null, and it is itself a random variable: ifH₀is true and the test assumptions hold,pis uniformly distributed on[0, 1]. That is exactly whyP(p \le \alpha \mid H_0) = \alpha— rejecting whenp ≤ αis what delivers the error rate you chose.Viji’s test. Two independent samples,
n₁ = n₂ = 30, pooleds ≈ 13: $$ t = \frac{\bar{x}_2 - \bar{x}_1}{s\sqrt{2/n}} = \frac{8}{13\sqrt{2/30}} = 2.383, \quad df = 58, \quad p = 0.020 $$ The same8-cup difference withn = 10per side givest = 1.376,df = 18,p = 0.186. Identical effect, identical spread, different evidence — because the standard error carries√n.What p is not. It is not
P(H₀ \mid data); getting that requires a prior and Bayes’ theorem. It is not the probability the result was “due to chance”. It is not1 −the probability of replication. Andp = 0.20is not evidence forH₀— absence of evidence is not evidence of absence, which is the Type II side of the Type I and Type II story.Why the threshold is arbitrary. Fisher offered
0.05as convenient and explicitly expected researchers to choose their own. Neyman and Pearson then built the fixed-α decision procedure the modern test inherits. The two frameworks answer different questions and were fused into one recipe afterwards — which is whypfeels like evidence andαfeels like a rule, and why they sit awkwardly together.What to report instead of a verdict. The effect size with a confidence interval, the exact
p(notp < 0.05), the sample size, and the α fixed in advance. Where many tests are run, control the family-wise error rate (Bonferroni) or the false discovery rate (Benjamini–Hochberg) — because with20independent tests atα = 0.05, the chance of at least one false alarm is1 − 0.95²⁰ ≈ 64%.
She draws the line. He measures the distance. Neither of them is allowed to do the other’s job afterwards.