It’s a Saturday evening. The sun has just gone. The kettle hisses. The lantern above the counter has been lit. Viji has had a good week, and he has two new things to show for it.
The first is a digital kitchen scale. He bought it on Monday to weigh out tea leaves, because the old spring scale’s needle wobbles and he never trusted it.
The second is a small customer card. He printed 40 of them: 4 questions, each scored from 1 to 5:
1. The tea tastes good. 2. I get served quickly. 3. The price is fair. 4. I would recommend this stall.
He handed them out all week, collected all 40, and averaged them this afternoon: 4.6 out of 5.
He is planning to raise the price of a glass of tea by ₹2. Customers this happy, he reasons, will not mind.
His cousin, as always, has a view: “4.6? Brother, you handed those cards out yourself, standing right there, with a free biscuit. I would have given you a 5 for the biscuit.”
Viji types back, “The card is very consistent. Everyone answers the same way.”
“That is exactly what worries me.”
Viji frowns at the phone. “Consistent is good,” he mutters. “Consistent is the whole point.”
The awning rustles.
A woman steps in under the awning: precise, composed, slate-blue shirt, hair pulled back tight. In one hand she holds a small wooden rubber stamp; in the other, a tin ink pad. Without a word she takes a sheet of paper from his counter and stamps it 5 times. The 5 marks land almost on top of each other.
She says, “I am Reliability. I am the promise that if you measure the same thing again, you get the same answer again. Same scale, same packet, same reading. Same customer, same card, same score. Look at my marks. Five times, one place.”
A man walks in behind her: tall, unhurried, saffron shirt, reading glasses on a cord. He is holding an old brass compass, lid open, and turning slowly on the spot until the needle settles.
He says, “I am Validity. I do not care whether you get the same answer twice. I care whether you are pointing at the right thing. Her stamp is very steady.” He glances at the sheet. “It is also a long way from where it should be.”
Viji looks at the sheet. The 5 stamps are tightly clustered, and they sit a hand’s width from the centre of the page.
They say, together, “Steady and correct are two different promises.”
Reliability speaks first
Reliability picks up a sealed packet of tea from under the counter. The label reads 250 g, from the wholesaler, who weighs on a certified scale.
She says, “Put it on your new scale.”
Viji does. The display reads 300 g.
“Take it off. Put it back.”
300 g. Again: 300 g. Twice more: 300 g, 300 g.
Reliability says, “Five readings, no spread at all. I love this scale. It is the most reliable instrument in your stall. Now take out the old one.”
He fetches the spring scale from the shelf and weighs the same packet 5 times. The needle lands on 232, 268, 244, 261, 245.
Viji adds them up in the notebook. “1,250. Average 250. The old one is right.”
Reliability says, “On average. But which of those 5 readings would you trust tomorrow morning, when you weigh one packet once and sell it? You never get to average. You get one reading, and on this scale one reading can be off by 18 grams in either direction. It is not reliable, and because it is not reliable, no single reading from it can be trusted. That is the part people forget.”
She sets her stamp down.
“Here is my claim to dignity. Nobody checks validity first. They check me, because if I fail, nothing else is worth checking. A measure that gives a different answer every time cannot be measuring anything in particular. In a questionnaire I come in three kinds: whether the same people give the same answers next week (test-retest), whether two observers agree when they score the same thing (inter-rater), and whether your questions agree with each other (internal consistency). That last one is the number everyone quotes: Cronbach’s alpha.”
Viji says, “What’s my card’s alpha?”
She runs the 40 cards through the calculator on his phone. “0.86. Above 0.7 is usually called acceptable. 0.86 is good. Your card is reliable.”
Viji says, “Then it’s working.”
Reliability says, “Then it is consistent. Ask him the other question.”
Validity speaks
Validity closes the compass lid and taps it on the counter.
He says, “Your new scale is reliable and wrong. It is consistently measuring something. It is just not measuring the weight of tea. It is measuring the weight of tea plus 50 grams. Very steadily.”
“Now your card. You told me it measures satisfaction. So here is my test. What should satisfied customers do?”
Viji says, “Come back.”
“Do you record who comes back?”
Viji does. His notebook marks regulars by name. He and Validity spend 10 minutes lining up each of the 40 card scores against how many times that customer returned in the 3 weeks since.
Validity reads the result. “The correlation between a customer’s card score and how often they came back is 0.08.”
Viji stares at it. “That’s almost nothing.”
“It is almost nothing. Satisfied customers should return more often. Your card cannot tell the ones who return from the ones who don’t. So whatever it measures, it is not satisfaction. My guess is that it measures how polite a person feels towards the owner who is standing in front of them, holding a biscuit. And it measures that very reliably.”
He lets that sit.
“This check has a name. Comparing a measure against an outcome it should predict is criterion validity. There are others. Content validity asks whether your questions cover the whole idea. Construct validity asks whether your score behaves the way the concept should: high where it should be high, unrelated to things it should be unrelated to. None of them is a single number you can look up, which is exactly why people skip me and quote her alpha instead.”
What just happened
Viji pours himself a glass of tea for the first time all evening and sits down on his own counter.
He says, slowly, “So the new scale and the card have the same problem. They’re both steady. They’re both pointing at the wrong thing.”
Both nod.
He says, “And the old spring scale has the opposite problem. It’s pointing roughly at the right thing on average, but it’s too shaky to trust any one reading.”
Validity says, “Now the asymmetry, which is the most important sentence tonight. A measure can be reliable without being valid. Your new scale proves it. But a measure cannot be valid without being reliable, because if the reading jumps around at random, then no single reading can be tracking the truth. She is necessary but not sufficient for me.”
Reliability says, “In other words, I am the floor. He is the building.”
Viji writes the grid in his notebook:
| Valid (points at the right thing) | Not valid | |
|---|---|---|
| Reliable (steady) | the goal | new scale · the biscuit card |
| Not reliable (shaky) | impossible for single readings | the worst case |
The same chat, in a chart
That picture is the same conversation, drawn. The four targets are the grid in Viji’s notebook: steady-and-centred, steady-and-off, shaky, and both. The two scatter plots are his card before and after, and they show why a higher alpha is not a better card. The biscuit card has the higher alpha and predicts nothing; the anonymous card has a slightly lower alpha and actually tracks who comes back.
One last warning before they leave
Reliability wipes her stamp on a cloth and caps the ink pad. She says, “One trap, and it is mine, and I am tired of it.”
“People see a high alpha and they stop thinking. So they try to push it higher. And the fastest way to push alpha higher is to ask the same question several times.”
She writes 4 questions on the back of one of Viji’s cards:
The tea tastes good. · The tea is nice. · I like the tea. · The tea is tasty.
“Give that to 40 people and your alpha will be about 0.97. Reviewers will call it excellent. It is not excellent. It is 1 question asked 4 times, and it only measures taste. It says nothing about speed, or price, or whether they would come back. You raised me by starving him.”
Validity adds, “Content validity. A satisfaction card with no question about price cannot tell you whether a price rise is safe, and that was your actual decision. Above about 0.95, alpha is more often a warning of redundancy than a sign of quality.”
Viji writes it down. Alpha is consistency, not truth. Too high is its own problem.
The bill
They left the way people leave a tea stall at night: Reliability capping her ink pad with a neat click, and Validity checking his compass once more under the lantern before stepping out into the dark.
Viji did three things that week. He put a 250 g packet on the new scale every morning and subtracted 50 from every reading, which made it both reliable and valid. He printed 60 new cards, kept all 4 questions, and rewrote question 2 the other way round, “I often wait too long to be served”, so that people had to read before ticking. And he stopped handing the cards out. He put a wooden box with a slot at the end of the counter, with no biscuit and nobody watching.
52 cards came back. The average was 3.9, not 4.6. The alpha was 0.81, a little lower. And the correlation with return visits was 0.54.
The price question, it turned out, was the lowest-scoring item on the card. He did not raise the price.
His cousin texted: “3.9? That’s worse, brother.”
Viji replied, “It’s lower. It’s not worse. The other one was a biscuit score.”
He turned to a fresh page and wrote the sentence he wanted to keep:
Reliability asks: do I get the same answer twice? Validity asks: is it the right answer? A steady wrong answer is still wrong.
For the math-curious
Cronbach’s alpha. For a scale of
kitems, with item variancesσ²ᵢand total-score varianceσ²ₜ: $$ \alpha = \frac{k}{k-1}\left(1 - \frac{\sum_{i=1}^{k} \sigma^2_i}{\sigma^2_t}\right) $$ When items move together, the total varies far more than the items do individually, and alpha rises towards1. Alpha also rises mechanically withk, so a long scale can look consistent even when its items are only weakly related. Alpha assumes every item measures the same single thing equally well. When that assumption is doubtful, McDonald’s omega or composite reliability from a factor model is the usual alternative.Reverse-worded items. Recode them before computing alpha: on a
1–5scale, recoded =6 − original. Forgetting this is one of the most common causes of a suspiciously low, or negative, alpha.Validity, in the vocabulary reviewers use.
- Content validity: do the items cover the full construct? Established by argument and expert judgement, not by a statistic.
- Criterion validity: does the score correlate with an outcome it should predict? Concurrent if measured at the same time, predictive if measured later, like Viji’s return visits.
- Construct validity: does the score behave as theory says? Convergent validity: it correlates with measures of the same idea. Discriminant validity: it does not correlate strongly with measures of different ideas. In structural equation modelling these are often assessed with AVE (average variance extracted, conventionally
≥ 0.5) and the Fornell–Larcker criterion or the HTMT ratio (conventionally below0.85or0.90).Why reliability caps validity. In classical test theory, the correlation between a measure
Xand any criterionYcannot exceed the square root of the product of their reliabilities: $$ r_{XY} \le \sqrt{r_{XX’} \cdot r_{YY’}} $$ A measure with reliability0.49cannot correlate above0.7with anything, however valid the idea behind it. That inequality is the formal version of necessary but not sufficient.
A steady answer tells you the instrument is not guessing. It does not tell you what the instrument is looking at.