Tests an intervention's effect when random assignment is impossible — using comparison groups, pretests and time series to rule out rival explanations one by one.
Research designQuantitative / causal-comparative
What it is
Quasi-experiments occupy the ground between correlational studies and true experiments. There is a genuine intervention or treatment — a new appraisal system, a minimum-wage change, a store redesign — but participants are not randomly assigned to it. Who gets treated is determined by policy, geography, timing or self-selection. The researcher's craft lies in adding design features — nonequivalent comparison groups, pretests, multiple time points — that make rival explanations progressively less plausible.
The workhorse forms are: the nonequivalent-groups pretest–posttest design (a treated unit and a comparison unit, both measured before and after — supporting difference-in-differences logic); the interrupted time series (many observations before and after an intervention, looking for a break in level or trend); and the one-group pretest–posttest design (weakest, as history and maturation remain uncontrolled). Regression-based matching of treated and untreated cases on observed covariates is a further strengthening tool.
Because groups may differ before treatment, every quasi-experimental report must argue explicitly against the classic threats to internal validity: selection (groups differed to begin with), history (something else happened at the same time), maturation, testing and instrumentation effects, and regression to the mean. A quasi-experiment without a threats discussion is just a before–after anecdote.
✓ When to use
A real intervention exists but random assignment is impossible, unethical or refused by the organization
Policy or natural experiments — a law, system rollout or restructuring applied to some units and not others
Evaluation of programmes already implemented, where only observational and archival data remain
Pilot rollouts where one site adopts first and comparable sites can serve as controls
Long archival outcome series exist, enabling interrupted time-series analysis around the intervention date
✗ When NOT to use
Random assignment is actually feasible — do not settle for a weaker design out of habit
No credible comparison group or pretest data can be obtained AND no time series exists — causal claims would be unsupported
The treated and comparison groups differ on the very factors that drive the outcome, and nothing measured can adjust for it
The intervention coincided exactly with another major event affecting the outcome (history confound you cannot separate)
You only need to describe or correlate — do not force causal framing where none is required
Typical research questions
RQ — Did the four-day workweek pilot in one business unit reduce absenteeism relative to comparable units that kept the five-day week?
RQ — Did mandatory POSH training, rolled out region by region, change harassment reporting rates?
RQ — Did the loyalty-programme relaunch increase repeat-purchase frequency, judged against 24 months of prior sales data?
RQ — Do employees who opted into hybrid work show different engagement trajectories than office-based peers, after matching on role and tenure?
Key characteristics
Purpose
Estimate an intervention's causal effect when assignment to treatment is not under the researcher's control
Typical data
Quantitative — outcome measures before and after treatment, often archival series; covariates for matching and adjustment
Researcher control
Partial: the treatment is real but assignment is not randomized; control comes from design features (comparison groups, pretests, time series)
Temporal aspect
Inherently longitudinal in logic — pre and post measurements, sometimes long observation series
Typical sample
Whatever the setting provides: treated and comparison units (sites, teams, regions) plus individuals within them; more pre/post time points beat more subjects for some threats
Variables & measurement
The independent variable is treatment exposure (treated vs comparison, or pre vs post), which you record rather than assign. The critical extra variables are pretest measures of the outcome and covariates describing how the groups differ — role mix, size, baseline performance — because these carry the entire burden of the selection-bias argument.
Measure the outcome identically, with the same instrument and procedure, in both groups and at every time point. Instrumentation changes (a new recording system mid-study) masquerade as treatment effects.
Collect the richest feasible pretest: multiple pre-intervention time points reveal whether groups were already trending apart
Document how treatment assignment actually happened — who chose, on what basis — since that is the selection mechanism you must argue about
Record concurrent events (policy changes, market shocks) as potential history confounds
Where possible add a non-equivalent dependent variable: an outcome that should NOT respond to the treatment but would respond to the confounds
Sampling approaches that fit
Sampling is usually of units and time points, not individuals: choose comparison sites/teams as similar as possible on structure, baseline outcome levels and trends
Matching — pair treated cases with untreated cases on key covariates (or propensity-score logic) to build a fairer comparison
Within units, census or systematic sampling of records is common (all employees, all transactions in the window)
For interrupted time series, plan enough pre- and post-intervention observations (dozens of points are far better than a handful)
Report why the comparison group is credible — that argument does the work randomization would have done
Data collection methods that fit
Archival and administrative records — HRIS, sales, safety and attendance data are the backbone of most quasi-experiments
Repeated surveys with identical instruments before and after the intervention in both groups
Organizational documentation of the intervention itself: what was implemented, where, when, with what fidelity
Field observation to verify the comparison group genuinely remained untreated (no informal spillover)
A timeline log of concurrent events for the threats-to-validity discussion
Appropriate statistical & analytical methods
The analytic theme is comparison with adjustment: compare treated and untreated groups on change in the outcome, adjusting for pre-existing differences; or model the outcome series and test for a break at the intervention. All links open the Statistical Methods Atlas.
Each link below opens the matching method page (opens the Statistical Methods Atlas) with assumptions, worked examples and reporting guidance.
However sophisticated the adjustment, state the identifying assumption openly — e.g. “absent the intervention, the two units would have followed parallel trends” — and show whatever evidence you have for it (parallel pre-trends, covariate balance after matching).
Quality criteria
Internal validity is the battleground — address each classic threat by name: selection, history, maturation, testing, instrumentation, regression to the mean, and attrition
Design strength hierarchy — one-group pretest–posttest < nonequivalent groups posttest-only < nonequivalent groups pretest–posttest < interrupted time series with comparison series; choose the strongest the setting allows
Construct validity — verify the intervention was implemented as described (fidelity), and outcomes measured consistently
Statistical conclusion validity — respect the clustered and autocorrelated structure of organizational data
External validity — effects estimated in one site under one selection regime may not transfer; replication across sites is the remedy
A worked mini-example
A company introduces a wellness programme in its Pune campus (2,100 employees) while the comparable Chennai campus (1,900 employees) continues as before. The researcher obtains 12 months of pre- and post-launch monthly absenteeism data for both campuses plus employee covariates. Pre-launch trends are statistically parallel — the key credibility check.
A difference-in-differences analysis (campus × period interaction) shows absenteeism falling 0.9 days/quarter more in Pune than in Chennai (95% CI [0.3, 1.5]). Rival explanations are examined: no concurrent policy differed between campuses; instrumentation was identical (same HRIS); a non-equivalent dependent variable (voluntary training hours) shows no parallel jump, weakening a general-morale-shock explanation. The report claims a “programme effect under the parallel-trends assumption”, not proven causation.
Common pitfalls
Treating a quasi-experiment as if it were randomized — no baseline analysis, no threats discussion
One-group before–after studies attributing any change to the intervention (history and maturation are always rival suspects)
Comparison groups chosen for convenience that differ on the outcome's main drivers
Regression to the mean — sites selected for treatment because they were performing worst will improve anyway
Ignoring autocorrelation in time-series outcomes, wildly overstating significance
Adjusting for post-treatment variables, which can absorb the very effect being estimated
Silent instrumentation changes (new measurement system introduced along with the intervention)
What to report
The specific quasi-experimental design used, named (e.g. nonequivalent-groups pretest–posttest; interrupted time series)
How units came to be treated — the actual selection mechanism — and why the comparison group is credible
Baseline equivalence and pre-trend evidence; matching or adjustment procedures
The identifying assumption stated in plain words, with supporting checks
Effect estimates with confidence intervals, from analyses that respect clustering and autocorrelation