Methodology Navigator
Navigator › Steps of the journey › From Analysis to Interpretation

From Analysis to Interpretation

Choosing the right analysis, reading results honestly — significance versus effect size — and writing conclusions your design can actually support.

Methodology step

Choosing the analysis: let the design decide

By this stage the hard choices are already made: your question type, design, variable roles and levels of measurement jointly determine a short list of appropriate analyses. Choosing the analysis first and bending the study around it is the classic beginner inversion. The practical selection logic:

The full decision tree — from data type to specific test, with assumptions and software steps — lives in the Statistical Methods Atlas's “Which test should I use?” helper on its home page.

Each link below opens the matching method page (opens the Statistical Methods Atlas).

Shapiro–Wilk Testcheck normality before choosing parametric vs nonparametric testsLevene's Testcheck homogeneity of variance across groupsIndependent-Samples t-Testthe canonical two-group comparison — a good anchor for the logicPearson Correlationthe canonical association measure — and the classic causation trapMultiple Regressionprediction with several predictors, each adjusted for the others

Statistical significance: what p actually says

A p-value is the probability of observing data at least this extreme if the null hypothesis were true. It is not the probability that the null is true, not the probability the finding replicates, and not a measure of importance. “p < .05” tells you the data are surprising under no-effect; it does not tell you the effect matters.

Two consequences follow. First, with a large sample, trivial effects become significant — n = 10,000 will flag a correlation of .03. Second, with a small sample, real effects go undetected — a nonsignificant result in an underpowered study is not evidence of absence. Always read p alongside the effect size and its confidence interval.

Effect size: how much, not just whether

Effect sizes express magnitude in interpretable units and are what practical decisions and meta-analyses run on. Report one, with a confidence interval, for every focal test.

Cohen's d (group differences)Difference between means in SD units. Conventional benchmarks: 0.2 small, 0.5 medium, 0.8 large — context always trumps benchmarks.
Correlation rDirection and strength of association: ±.10 small, ±.30 medium, ±.50 large by convention. r² is variance shared.
R² / adjusted R² (regression)Proportion of outcome variance the model accounts for; adjusted R² penalizes predictor count.
η² / partial η² (ANOVA)Proportion of variance attributable to a factor: .01 small, .06 medium, .14 large by convention.
Odds ratio (logistic models)Multiplicative change in odds per unit of the predictor; OR = 1 means no association.

Causal language discipline

The verbs in your conclusions must be earned by the design, not by the statistics. The same β = .40 supports very different sentences depending on how the data arose.

Randomized experimentCausal verbs earned: “the training increased performance”, “reduced”, “led to” — within the studied setting and population.
Quasi-experimentConditional causal language: “consistent with a programme effect, assuming parallel trends”; name the threats that remain.
Correlational / survey studyAssociational verbs only: “is associated with”, “predicts” (statistically), “is related to”. Not “influences”, “drives”, “impacts”, “enhances”.
Qualitative studyInterpretive claims: “participants experienced X as…”, “the accounts suggest a mechanism whereby…” — offered as transferable insight, not population estimates.

Reporting standards

Interpretation: closing the loop

Interpretation returns the statistics to the research question: What did we ask? What did the data answer, at what magnitude and with what uncertainty? What rival explanations survive the design? What should a reader — scholar or practitioner — now do differently? A disciplined interpretation section connects each finding back to its hypothesis, sizes it against prior literature, owns the limitations specifically (not the ritual “small sample, future research needed”), and states scope conditions: the population, setting and time your claims cover.

Then stop. The most common interpretive failure is not statistical error but overreach — conclusions that quietly outrun the design. If you began with a sharp question and a fitting design, the honest conclusion writes itself.

Common pitfalls

Related pages

Correlational Researchwhere associational language discipline matters mostExperimental Researchthe design that earns causal verbsQuasi-Experimental Researchconditional causal claims and threats-to-validity reportingResearch Questions & Hypothesesinterpretation answers the question you wrote at the start
← Data Collection Methods