Choosing the right analysis, reading results honestly — significance versus effect size — and writing conclusions your design can actually support.
By this stage the hard choices are already made: your question type, design, variable roles and levels of measurement jointly determine a short list of appropriate analyses. Choosing the analysis first and bending the study around it is the classic beginner inversion. The practical selection logic:
Each link below opens the matching method page (opens the Statistical Methods Atlas).
A p-value is the probability of observing data at least this extreme if the null hypothesis were true. It is not the probability that the null is true, not the probability the finding replicates, and not a measure of importance. “p < .05” tells you the data are surprising under no-effect; it does not tell you the effect matters.
Two consequences follow. First, with a large sample, trivial effects become significant — n = 10,000 will flag a correlation of .03. Second, with a small sample, real effects go undetected — a nonsignificant result in an underpowered study is not evidence of absence. Always read p alongside the effect size and its confidence interval.
Effect sizes express magnitude in interpretable units and are what practical decisions and meta-analyses run on. Report one, with a confidence interval, for every focal test.
| Cohen's d (group differences) | Difference between means in SD units. Conventional benchmarks: 0.2 small, 0.5 medium, 0.8 large — context always trumps benchmarks. |
|---|---|
| Correlation r | Direction and strength of association: ±.10 small, ±.30 medium, ±.50 large by convention. r² is variance shared. |
| R² / adjusted R² (regression) | Proportion of outcome variance the model accounts for; adjusted R² penalizes predictor count. |
| η² / partial η² (ANOVA) | Proportion of variance attributable to a factor: .01 small, .06 medium, .14 large by convention. |
| Odds ratio (logistic models) | Multiplicative change in odds per unit of the predictor; OR = 1 means no association. |
The verbs in your conclusions must be earned by the design, not by the statistics. The same β = .40 supports very different sentences depending on how the data arose.
| Randomized experiment | Causal verbs earned: “the training increased performance”, “reduced”, “led to” — within the studied setting and population. |
|---|---|
| Quasi-experiment | Conditional causal language: “consistent with a programme effect, assuming parallel trends”; name the threats that remain. |
| Correlational / survey study | Associational verbs only: “is associated with”, “predicts” (statistically), “is related to”. Not “influences”, “drives”, “impacts”, “enhances”. |
| Qualitative study | Interpretive claims: “participants experienced X as…”, “the accounts suggest a mechanism whereby…” — offered as transferable insight, not population estimates. |
Interpretation returns the statistics to the research question: What did we ask? What did the data answer, at what magnitude and with what uncertainty? What rival explanations survive the design? What should a reader — scholar or practitioner — now do differently? A disciplined interpretation section connects each finding back to its hypothesis, sizes it against prior literature, owns the limitations specifically (not the ritual “small sample, future research needed”), and states scope conditions: the population, setting and time your claims cover.
Then stop. The most common interpretive failure is not statistical error but overreach — conclusions that quietly outrun the design. If you began with a sharp question and a fitting design, the honest conclusion writes itself.