All posts Research

Probability and Non-probability Sampling Walk Into a Tea Stall

Viji asked the 40 people standing in front of him and got a clear answer: bun-butter. Then somebody drew 60 numbered tokens out of a drum and the answer changed completely. Two strangers explain why one of those samples has a margin of error and the other one cannot have one.

Vijayakumar P — Yuvijen Vijayakumar P 13 min read
Stylised illustration of a busy tea stall in daylight, a crowd of many small customer figures spread across the frame, each wearing a small numbered token. On the left a man in burnt orange has set down two brass posts with a rope cordon looped around the handful of people nearest the counter, and holds a clipboard. On the right a woman in deep teal turns a brass lottery drum, and scattered evenly across the whole crowd, far beyond the rope, a dozen figures are ringed in teal to show who was drawn. Viji stands behind the counter between them.

It’s a Saturday morning. The kettle hisses. The fan turns slowly. Viji is adding one item to the menu, and he has been arguing with himself about it for a week.

Three candidates: vada, boiled eggs, bun-butter. One fryer, one new item, and whichever he picks he is committed for the season.

So last Sunday he did the sensible thing and asked. He stood at the counter from 8 to 11 in the morning and asked every customer who came. 40 people answered.

Bun-butter 62%. Vada 23%. Eggs 15%.

That is not close. He has the supplier on the phone this afternoon.

His cousin, who ruins Saturdays for a living: “You asked Sunday morning, brother. Who comes on Sunday morning?”

Viji types back, “Customers.”

“Families. Grandfathers. Children. Not the 7 a.m. Tuesday crowd who are your actual business. You asked the ones who were standing in front of you.”

Viji puts the phone down and mutters, “Everybody is standing in front of me eventually.”

The awning rustles.

A man steps in under the awning — quick, practical, sleeves rolled, burnt orange shirt, a clipboard under one arm. In each hand he carries a small brass post, and looped between them a short length of red rope. He sets the posts down on either side of the counter and lets the rope hang across, penning in the six or seven people nearest the front.

He says, “I am Non-probability Sampling. I ask whoever is inside the rope. That is my whole method, and I want to be honest about it: I am fast, I am cheap, I need no list of anybody, and I gave you an answer in three hours on a Sunday.”

A woman walks in behind him — unhurried, deep teal shirt, carrying a brass lottery drum on a little crank handle. Inside it, tokens rattle. She sets it on the counter and turns the handle twice.

She says, “I am Probability Sampling. Before I ask anybody anything, I need one thing he never needs: a list. Every customer on it, each with a number. Then the drum decides, not me, and not who happens to be standing near the tea.”

Viji says, “I have a list. Sort of.” He pulls up the till records — 2,400 regular customers, each with a bill history.

She smiles. “Then you have what researchers call a sampling frame, and you are luckier than most people who ask me for help.”

Probability Sampling speaks first

She cranks the drum and draws out a token. 1,184. Then another. 76. Then another.

She says, “Sixty tokens out of 2,400. Here is the sentence that defines me, and it is worth memorising: every customer on that list has a known, non-zero chance of being picked. In this case, 60 in 2,400 — 1 in 40. Exactly the same chance for the Tuesday-morning worker as for the Sunday grandfather.”

Viji says, “Does it matter that it’s the same chance? What if it isn’t?”

“It does not have to be equal. It has to be known. If I deliberately give one group a better chance, I can correct for it afterwards by weighting each answer by 1 over its chance of selection. What I cannot correct is a chance I do not know — and that is his entire situation.”

She lays four tokens on the counter.

“I come in four common shapes. Simple random is what the drum just did — pure lottery. Systematic is cheaper: 2,400 ÷ 60 = 40, so pick a random start and take every 40th customer through the door. Stratified is the clever one: split your list into groups first — morning rush, afternoon, evening families — then draw randomly inside each group, in proportion. Cluster is the cheap one for scattered populations: pick 6 days at random and ask everybody who comes on those days.”

Viji ran the drum version that week. 60 customers, drawn by number, tracked down over four days.

Vada 54%. Bun-butter 27%. Eggs 19%.

He stared at it for a long time.

“It flipped.”

“It flipped because his rope was tied around a Sunday. Your morning-rush workers are about 70% of your trade and almost none of Sunday’s queue. They want something filling that survives a scooter ride. The grandfathers wanted bun-butter, and they were the only people you asked.”

Non-probability Sampling speaks

He does not argue. He coils the rope up and sits down.

He says, “She is right, and I want to be precise about what I got wrong. It was not that my 40 people lied. It was that the chance of a Tuesday worker landing inside my rope on a Sunday morning was zero — and I had no way of knowing that from inside the data. My 62% is a perfectly accurate description of 40 people who are not your customer base.”

He holds up the clipboard.

“I also come in four shapes, and you should know them, because you will use every one of them eventually. Convenience is what I just did — whoever is in front of me. Purposive is deliberate: pick the 12 customers who have come here for 10 years, because you want depth, not a headcount. Quota fills targets — 42 workers, 9 afternoon, 9 evening — but I choose who fills them. Snowball is for people you cannot see: you have never met the night-shift drivers who pass at 3 a.m., so you ask the one you know to introduce another, and he introduces a third.”

He says, “Here is my claim to dignity, and it is not a small one. For the night-shift drivers there is no list. She cannot draw a token for a person who appears on no register anywhere. Snowball is not my lazy option there — it is the only door in the building. For hidden populations, for early exploration, for the question what should I even be asking about, and for every qualitative study where depth is the point, I am the correct tool and she is the wrong one.”

He taps the clipboard.

“What I cannot do — ever — is hand you a margin of error. She can, because she knows the chance each person had of being chosen. I do not know mine, so no honest number comes out of my data with a ± in front of it, and nothing I produce generalises to your 2,400.”

What just happened

Viji pours a glass of tea, sits on his own counter, and writes the grid out.

ProbabilityNon-probability
Selection bychance, from a framejudgement, access, convenience
Every unit’s chanceknown, non-zerounknown, often zero
Needs a list?yesno
Margin of errorcomputablenot computable
Generalises to population?yes, with stated uncertaintyno
Cost and speedslow, expensivefast, cheap
Shapessimple random, systematic, stratified, clusterconvenience, purposive, quota, snowball
Right whenyou need a number for everyoneyou need depth, or there is no list

Probability Sampling adds one row underneath.

“And the shape you pick changes your precision even at the same n = 60.”

Simple random: ± 6.4 points. Stratified: ± 4.8 points. Cluster (6 days): ± 8.9 points.

She says, “Stratifying beats the lottery because you have forced the morning-rush proportion to be exactly right instead of leaving it to luck. Clustering loses because six days is really six groups of similar people — the 60 answers carry less information than 60 independent ones. Same sample size. Three different amounts of knowledge.”

Through the awning, the man with the brass ladle from a few months back raises his glass without getting up. He says, “I told you to stir the pot. These two are arguing about how to stir.”

The same chat, in a chart

Three-panel chart on pale apricot: Panel I shows a crowd of 2,400 small dots, with a red rope cordon enclosing a small cluster of 40 at one edge labelled convenience, chance unknown or zero, and beside it the same crowd with 60 dots ringed in teal scattered evenly throughout, labelled simple random, chance 1 in 40 for everyone. Panel II shows two bar charts side by side of the three menu options: the Sunday convenience sample of 40 giving bun-butter 62 percent, vada 23, eggs 15, and the probability sample of 60 giving vada 54 percent, bun-butter 27, eggs 19, with an arrow noting that the winner flips. Panel III is a small cartoon of Probability Sampling in deep teal turning a brass lottery drum and Non-probability Sampling in burnt orange carrying two brass posts with a red rope slung between them.

That picture is the same conversation, drawn. The first panel is the difference in one image: a rope around whoever was near, against tokens scattered across everyone. The second panel is what that difference cost — the same question, the same stall, two weeks apart, and a different winner.

One last warning before they leave

Probability Sampling stops the drum.

She says, “Three traps. The first one is the most common mistake in student research anywhere.”

“One. Quota is not stratified, and it is not mine. They look identical on paper — the same groups, the same proportions, 42 and 9 and 9. The difference is the step nobody writes down: stratified draws randomly inside each group, quota lets the interviewer fill each group with whoever is willing. Quota has the shape of a probability design and none of its licence. If your method section says quota and your results section says margin of error, a reviewer will stop reading there.”

“Two. Systematic sampling has one specific enemy: a rhythm. Every 40th customer is fine until 40 happens to align with something in your stall’s cycle — a shift change, a bus arrival, a market day. Then your sample quietly fills with one kind of person. Check that your interval does not match any natural period, or randomise the start each day.”

He picks up his rope and adds the last one himself.

“Three, and this one is mine to give, because it turns her into me. A perfect random draw of 60 names means nothing if only 20 of them answer. The 40 who ignored you are not a random 40 — busy people, unhappy people, and people who dislike surveys are missing in a pattern. That is nonresponse bias, and past roughly 20% or 30% missing, your probability sample has quietly become one of mine, with all of my limits and none of my honesty about them. Chase the ones who did not answer. They are the expensive half.”

Viji writes all three down. Quota is not stratified. Watch the rhythm. Chase the missing.

The bill

They left the way people leave a tea stall on a working Saturday. Non-probability Sampling gathered his two brass posts under one arm, coiled the rope, and went off at speed to ask somebody something. Probability Sampling took longer, because she stopped to write the 60 token numbers into the back of Viji’s notebook first.

Viji ordered the vada.

He also did something he would not have thought of a month earlier. For the night-shift drivers — the ones who pass at 3 a.m., who are on no list of his, who have never once been inside anybody’s rope — he asked the one driver he knows to mention it to the others. Four of them came by that week. Three wanted eggs, because eggs keep.

So he ordered eggs too, in small numbers, for the night window only. That decision came from a snowball sample of 4 people and it is not generalisable to anybody, and it did not need to be. It was a question about 4 drivers.

His cousin: “So I was right about Sunday.”

Viji typed back, “You were right about Sunday. You were wrong that asking people is useless. It depends entirely on who gets a chance to be asked.”

He turned to a fresh page and wrote the sentence he wanted to keep:

A sample is only a shortcut to the population if everyone in the population could have been in it. Otherwise it is just a group of people you spoke to.


For the math-curious

The formal line. A design is a probability sample if every unit in the frame has a known inclusion probability πᵢ > 0. That single condition is what licenses everything downstream: unbiased estimation, standard errors, confidence intervals, and generalisation to the frame. Unequal πᵢ is fine — the Horvitz–Thompson estimator weights each unit by 1/πᵢ: $$ \hat{T} = \sum_{i \in S} \frac{y_i}{\pi_i} $$ Non-probability designs have πᵢ unknown, and for some units πᵢ = 0, which no weight can repair.

Margin of error for a proportion (simple random): $$ \text{SE}(\hat{p}) = \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} \qquad \text{MoE}_{95%} = 1.96 \times \text{SE} $$ Viji’s numbers: p̂ = 0.54, n = 60, SE = 0.064, so about ± 6.4 percentage points at 1 SE and ± 12.6 at 95%. For a sample that is a noticeable share of a finite population, multiply by the finite population correction from the Population and Sample post.

Why stratification wins. Stratified sampling removes between-stratum variance from the error term; only within-stratum variance remains. Gains are largest when strata are internally similar and differ from each other, which is exactly Viji’s morning-rush and Sunday-family split. Proportional allocation sets nₕ ∝ Nₕ; Neyman allocation also samples more heavily where the stratum is more variable.

Why clustering loses. The design effect is $$ \text{deff} = 1 + (m - 1),\text{ICC} $$ for clusters of size m with intraclass correlation ICC. People sharing a day resemble each other, so ICC > 0 and the variance inflates. The effective sample size is n / deff — Viji’s cluster margin of 8.9 against a simple-random 6.4 implies deff ≈ 1.9, so his 60 clustered answers carry roughly the information of 31 independent ones. Clusters are chosen for cost, not precision.

Nonresponse. With response rate r and a difference d in the outcome between responders and non-responders, the bias is about (1 - r) × d. Note it does not shrink with n: surveying 10,000 people at a 20% response rate buys precision around a biased number, which is the same failure the Population and Sample post found in the 1936 Literary Digest poll.

Repairing non-probability samples. Post-stratification, raking, propensity weighting and MRP (multilevel regression and post-stratification) can adjust a non-probability sample towards known population totals. All of them assume that, within the adjustment cells, responders and non-responders are alike — an assumption the data cannot check.

Everybody was standing in front of him eventually. The question was who was standing there on the day he asked.

Stay in the loop

Follow Yuvijen on LinkedIn.

New posts, research notes, and analytics tips — straight to your LinkedIn feed.

Follow on LinkedIn

linkedin.com/company/yuvijen · no signup needed

Free newsletter

One new explainer a week, in plain language

Statistics and analytics concepts explained simply — new posts, new tools, no spam.

No spam · unsubscribe anytime