All posts Analytics

p-value and Alpha Walk Into a Tea Stall

Viji's new kettle looks like it sold 8 more cups a day, and the test returns p = 0.020. His cousin asks what would have happened if he had wanted α = 0.01 instead. Two strangers explain why that question has to be answered before the data arrives, not after.

Vijayakumar P — Yuvijen Vijayakumar P 11 min read
Stylised illustration of the road outside a tea stall. A woman in slate blue kneels and snaps a chalk line across the ground before anything is thrown, her reel still in hand. A man in burnt orange stands further along with a measuring tape, waiting to measure where a stone has landed. Viji watches from behind his counter. The chalk line is labelled α = 0.05 and drawn before; the tape is labelled p = 0.020 and measured after.

It’s a Tuesday morning. The kettle hisses — the new kettle, which boils in half the time and cost him ₹3,100.

Viji bought it six weeks ago on the theory that a faster kettle means a shorter queue, and a shorter queue means nobody walks off to the stall near the bus stand. He has 30 days of customer counts from before it arrived and 30 from after.

Before: mean 87 customers a day. After: mean 95. SD about 13 in both stretches.

He runs the two-sample t-test he learned from the last two visitors. The output is one line:

t = 2.383, df = 58, p = 0.020

He knows the rule. 0.020 is less than 0.05, so he rejects the null and concludes the kettle did something.

His cousin, who has developed an unfortunate habit of asking the second question: “And if you had wanted to be strict, brother? If you had used 0.01 instead of 0.05?”

Viji looks at the number. 0.020 is not less than 0.01.

He types back, “Then I wouldn’t reject.”

“So which is it?”

Viji puts the phone down. “It’s the same kettle,” he says, to nobody. “The same 60 days. The same 0.020. How can the answer depend on a number I get to pick?”

The awning rustles.

A woman steps in under the awning — deliberate, unhurried, slate-blue shirt, with a builder’s chalk line reel clipped to her belt. She does not sit. She walks out to the road in front of the stall, kneels, pulls the string taut across the dust, and snaps it. A crisp blue line appears on the ground.

She says, “I am Alpha. I am that line. I am drawn before anything is thrown.”

A man walks in behind her — quick, cheerful, burnt-orange shirt, a steel measuring tape hooked on one finger. He does not look at the line at all. He is looking down the road, at where something has already landed.

He says, “I am the p-value. I measure after. Somebody throws, the stone lands, and I walk out with my tape and report how far from the mark it fell. That is the whole of my job, and I have no opinion at all about whether the distance is good enough.”

Viji looks between them. “But the two of you together decide whether my kettle worked.”

They reply, together, “Only in that order.”

Alpha speaks first

Alpha stands up and dusts off her hands. The blue line sits in the road, plain and permanent.

She says, “Here is what I actually am, and it is not what most people think. I am not a fact about your data. I have never seen your data. I am a policy — the error rate you have decided in advance that you are willing to live with.”

She says, “Set me at 0.05 and you are saying: if the kettle had done nothing at all, I accept a 5% chance of being fooled into saying it did. That is a statement about the procedure, made before the first day was recorded. It stays true whether your p comes out at 0.001 or 0.9.”

Viji says, “Then why can I choose it?”

“Because how much of that mistake you can afford is a question about your stall, not about your statistics. A 5% false-alarm rate on a kettle is nothing — you buy a kettle, and if it turns out to be useless, you own a slightly better kettle. If you were deciding whether to shut the stall and move to another road, you would want me much smaller, because that mistake costs you your living.”

She adds, with some emphasis, “And this is my one rule. I am drawn before the throw. Once the stone is in the air I do not move, and I am not open to negotiation. The moment I am chosen after seeing where it landed, I stop being an error rate and become a decoration.”

The p-value speaks

The p-value unhooks his tape and lets it run out along the road.

He says, “Now me. Everyone says my name and almost nobody says what I measure, so here it is, carefully.”

He says, “I assume the null is true. Not that it is — I assume it, for the sake of measuring. I imagine the kettle did nothing, that your 95 and your 87 differ only because 60 particular days fell the way they did. Then I ask one question: in that imaginary world, how often would I see a gap at least as big as 8 cups?”

“The answer is 0.020. Two times in a hundred.”

Viji says, “So there’s a 2% chance the kettle did nothing?”

“No. And that sentence is the single most common mistake in applied statistics, so let me be exact. I am not the probability that the null is true. I am the probability of data like yours, given that the null is true. Those are different questions, and swapping them is like confusing how often clouds appear when it rains with how often it rains when there are clouds.”

He walks the tape back in.

“Second thing about me, and it is the one that surprises people. I am not a measure of how big the effect is. Take your same 8-cup difference, and your same 13-cup spread, and collect only 10 days on each side instead of 30.”

He writes it in the dust: t = 1.376, p = 0.186.

Viji stares. “The same 8 cups. The same kettle. And now p is 0.186?”

“The same gap, measured with fewer days. I am part evidence and part sample size, and I never tell you which part is which. That is why a tiny effect in a huge study can post a thrilling p, and a large effect in a small study can post a dull one.”

He points at the chalk line. “Her line does not move when n changes. Mine does. We are built out of different material.”

What just happened

Viji pours a glass of tea, sits on his own counter, and writes the two of them out.

Alpha (α)p-value
Chosen or computed?chosen, by youcomputed, from the data
When?before the dataafter the data
What is it about?the procedurethis one sample
Depends on n?noyes
Changes if the data changes?noyes
Meaningerror rate you acceptP(data this extreme given H₀)
The decisionreject if p ≤ α

He looks at the last row for a long time.

He says, “So the cousin’s question was not really a question about statistics. He asked what I would have concluded at α = 0.01. But I had already chosen 0.05, weeks ago, before the kettle arrived. The honest answer is that 0.01 was never my line.”

Alpha says, “Correct. And if you had chosen 0.01 back then, you would have chosen it for a reason — a more expensive mistake — and you would be living with not rejecting today, and that would also be honest.”

The p-value adds, “What you must not do is look at my 0.020, notice it clears one bar and not the other, and then decide which bar you were aiming at.”

The same chat, in a chart

Animated three-panel figure. Panel one: a null distribution is drawn and a chalk line is snapped across it at the 5 percent mark, labelled alpha equals 0.05, chosen before the data. Panel two: the observed result lands at t equals 2.383 and a measuring tape runs out to shade the tail beyond it, labelled p equals 0.020, measured after, with the note that p is less than alpha so the null is rejected. An inset shows the same 8-cup difference measured over only 10 days a side, where p rises to 0.186 while the chalk line has not moved. Panel three: Alpha holding a chalk line reel and the p-value holding a steel measuring tape.

That picture is the same conversation, drawn. The first panel is the line, snapped before anyone looks. The second is the tape, run out after the stone lands. The inset is the part worth remembering: the stone travelled exactly as far, the line never moved, and the number changed anyway, because fewer days were used to measure it.

One last warning before they leave

The p-value clips his tape shut.

He says, “Three traps, and every one of them is somebody moving her line after the fact.”

“One. Do not choose α after seeing p. It has a name — and if you go looking for a threshold your result happens to clear, your stated error rate is fiction. The 0.05 has to be older than the data.”

“Two. Do not keep collecting until I behave. If you test at 30 days, then add 10 more because I sat at 0.08, then check again, you have taken several shots at the line while only admitting to one. Under repeated peeking, a nominal 5% false-alarm rate can rise past 20%. Fix the sample size in advance, or use a design that is built for interim looks.”

Alpha says, “Three, and this one is mine. 0.05 is a convention, not a law. Ronald Fisher suggested it as a convenient rule of thumb and said plainly that a worker should set his own level for his own purpose. Treating p = 0.049 as a discovery and p = 0.051 as nothing is treating a hand-drawn line as a cliff edge.”

She adds, “So report me, report him, and report the thing neither of us measures: the size of the difference, with a confidence interval around it. 8 more cups a day, somewhere between about 1 and 15, is a sentence your supplier can act on. p < 0.05 is not.”

Viji writes all three down. Draw the line first. Never move it. And say how big the difference was.

The bill

They left the way people leave a tea stall on a working morning. Alpha wound her chalk line back onto its reel and stepped over the blue mark on her way out without looking at it. The p-value coiled his tape, tapped it twice against his palm, and followed her.

Viji kept the kettle. He wrote the decision in the back of his notebook the way they told him to:

α = 0.05, fixed before the kettle arrived. Observed +8 cups a day, 95% CI roughly 1 to 15. p = 0.020. Rejected. Kettle stays.

Then he added the line he expected to need more often than the rest of it: p is not the chance I am wrong. It is how surprising my days would be if the kettle had done nothing at all.

His cousin, inevitably: “And if I still think 0.01 is the right bar?”

Viji typed back, “Then say so before the next thing I buy, and I’ll test it at 0.01. You don’t get to pick after you’ve seen the number.”

The cousin’s reply took some time. “Fair.”


For the math-curious

The definition. For an observed statistic T = t and a null hypothesis H₀, the two-sided p-value is $$ p = P\big(|T| \ge |t| ;\big|; H_0\big) $$ It is a tail probability under the null, and it is itself a random variable: if H₀ is true and the test assumptions hold, p is uniformly distributed on [0, 1]. That is exactly why P(p \le \alpha \mid H_0) = \alpha — rejecting when p ≤ α is what delivers the error rate you chose.

Viji’s test. Two independent samples, n₁ = n₂ = 30, pooled s ≈ 13: $$ t = \frac{\bar{x}_2 - \bar{x}_1}{s\sqrt{2/n}} = \frac{8}{13\sqrt{2/30}} = 2.383, \quad df = 58, \quad p = 0.020 $$ The same 8-cup difference with n = 10 per side gives t = 1.376, df = 18, p = 0.186. Identical effect, identical spread, different evidence — because the standard error carries √n.

What p is not. It is not P(H₀ \mid data); getting that requires a prior and Bayes’ theorem. It is not the probability the result was “due to chance”. It is not 1 − the probability of replication. And p = 0.20 is not evidence for H₀ — absence of evidence is not evidence of absence, which is the Type II side of the Type I and Type II story.

Why the threshold is arbitrary. Fisher offered 0.05 as convenient and explicitly expected researchers to choose their own. Neyman and Pearson then built the fixed-α decision procedure the modern test inherits. The two frameworks answer different questions and were fused into one recipe afterwards — which is why p feels like evidence and α feels like a rule, and why they sit awkwardly together.

What to report instead of a verdict. The effect size with a confidence interval, the exact p (not p < 0.05), the sample size, and the α fixed in advance. Where many tests are run, control the family-wise error rate (Bonferroni) or the false discovery rate (Benjamini–Hochberg) — because with 20 independent tests at α = 0.05, the chance of at least one false alarm is 1 − 0.95²⁰ ≈ 64%.

She draws the line. He measures the distance. Neither of them is allowed to do the other’s job afterwards.

Stay in the loop

Follow Yuvijen on LinkedIn.

New posts, research notes, and analytics tips — straight to your LinkedIn feed.

Follow on LinkedIn

linkedin.com/company/yuvijen · no signup needed

Free newsletter

One new explainer a week, in plain language

Statistics and analytics concepts explained simply — new posts, new tools, no spam.

No spam · unsubscribe anytime