Manipulates the independent variable and randomly assigns participants to conditions — the strongest design for establishing cause and effect.
Research designQuantitative / causal
What it is
A true experiment has three defining ingredients: the researcher manipulates the independent variable (some participants get the treatment, others a control or alternative), randomly assigns participants to those conditions, and controls extraneous influences. Random assignment is the crucial one — it makes the groups equivalent in expectation on every characteristic, measured or not, so a post-treatment difference in the outcome can be attributed to the treatment rather than to pre-existing differences.
Experiments in management research take several forms: laboratory experiments (scenario studies with students or online panels, high control, lower realism), field experiments (a real training programme randomly rolled out across branches — including the randomized controlled trial, or RCT), and survey experiments (randomizing question vignettes within a questionnaire). Between-subjects designs compare different people across conditions; within-subjects (repeated measures) designs expose the same people to all conditions and control individual differences at the cost of order effects.
The price of causal clarity is constraint. Many interesting causes cannot be manipulated ethically or practically, artificial settings can limit generalizability, and demand effects — participants guessing the hypothesis — can distort behaviour. A well-run experiment therefore pairs random assignment with manipulation checks, blinding where feasible, and honest discussion of how far the setting travels.
✓ When to use
The research question is explicitly causal — does X change Y? — and X can be manipulated
Random assignment to conditions is feasible and ethical
You need evidence strong enough to justify an intervention, policy or design decision
A correlational finding needs a causal test (e.g. does autonomy actually raise engagement, or do engaged people seek autonomy?)
Mechanisms need isolating — factorial designs can separate the effect of message framing from message source
✗ When NOT to use
The cause cannot be manipulated (gender, personality, firm size) or must not be (harmful stressors, deception beyond ethical bounds)
Random assignment is impossible in the setting — consider a quasi-experiment instead
The behaviour of interest cannot be evoked meaningfully in a controlled setting, and a field experiment is out of reach
The sample available is too small to detect a realistic effect size — an underpowered experiment mostly produces noise
Understanding meanings and processes is the goal — experiments quantify effects; they do not explain lived experience
Typical research questions
RQ — Does a structured onboarding programme, versus standard onboarding, increase new-hire retention at 6 months?
RQ — Does displaying scarcity cues (“only 3 left”) increase online purchase completion compared with identical pages without them?
RQ — Do resume-screening decisions differ when identical resumes carry male versus female names?
RQ — Does mindfulness training reduce emotional exhaustion in call-center agents relative to a waitlist control?
Key characteristics
Purpose
Establish causal effects: does manipulating X change Y, by how much, and under what conditions?
High: manipulation of the IV, random assignment, standardized procedures, manipulation checks
Temporal aspect
Prospective by construction — cause is imposed before the effect is measured; from one session to multi-month field trials
Typical sample
Sized by power analysis; typically dozens per condition in the lab, larger for field experiments with noisy outcomes
Variables & measurement
The independent variable is defined by the manipulation, so design effort goes into making conditions differ only in the intended ingredient: treatment vs control (ideally an active control that matches time and attention), or multiple levels and factors. Always include a manipulation check — a measure showing participants actually experienced the intended difference.
The dependent variable should be sensitive, reliable and as objective as the setting allows: behaviour and choices over self-reported intentions where possible. Decide covariates (e.g. a baseline measure of the outcome) in advance — they add precision but must be measured before randomization. Randomize order of materials in within-subjects designs to neutralize learning and fatigue effects.
Pilot the manipulation for strength and believability before the main study
Keep everything except the manipulated ingredient identical across conditions — instructions, timing, materials
Pre-register hypotheses, conditions, exclusions and the analysis plan where possible
Blind participants (and, where feasible, experimenters and raters) to condition and hypothesis
Sampling approaches that fit
Random assignment, not random sampling, is what secures internal validity — the two are routinely confused
Convenience samples (students, online panels such as Prolific/MTurk) are common and acceptable for theory-testing; discuss generalizability
Probability or organizationally complete samples strengthen field experiments aimed at policy conclusions
Determine n per condition by a priori power analysis (expected effect size, alpha, desired power — conventionally .80)
Plan for attrition in multi-session designs and report it by condition — differential dropout threatens the randomization
Data collection methods that fit
Laboratory or online sessions with standardized scripts, stimuli and timing
Field implementation with fidelity monitoring — did every branch actually deliver the assigned training?
Behavioural traces and system data as outcomes (clicks, purchases, error rates, retention) — resistant to demand effects
Validated scales for psychological outcomes, administered identically across conditions
Manipulation-check and attention-check items, plus a funneled debriefing to probe hypothesis guessing
Appropriate statistical & analytical methods
The analysis mirrors the design: compare conditions on the outcome, with the technique chosen by the number of groups, the assignment structure (between vs within subjects), and the outcome's measurement level. All links open the Statistical Methods Atlas.
Each link below opens the matching method page (opens the Statistical Methods Atlas) with assumptions, worked examples and reporting guidance.
Report effect sizes (Cohen's d, η²) with confidence intervals alongside p-values — a significant but tiny effect and a large one demand different conclusions. Analyse participants in the condition they were assigned to (intention-to-treat logic) in field experiments with imperfect compliance.
Quality criteria
Internal validity — the design's strong suit, protected by random assignment, control conditions, blinding and equal treatment of groups; threatened by differential attrition, contamination between conditions and demand effects
Construct validity of the manipulation — manipulation checks confirm the IV was experienced as intended, not something else (e.g. scarcity cue read as low quality)
External validity — artificial tasks, volunteer samples and short time frames limit generalization; field replication is the strongest answer
Ethics — informed consent, minimal deception with debriefing, and fair access to beneficial treatments (e.g. waitlist controls receive training later)
A worked mini-example
An HR researcher tests whether structured interviews reduce hiring bias. 120 practicing managers recruited through an executive program are randomly assigned to evaluate the same four candidate videos using either a structured scoring rubric (n = 60) or their usual unstructured judgment (n = 60). Order of candidates is counterbalanced; a manipulation check confirms rubric use.
Ratings of equally qualified male and female candidates differ by d = 0.42 in the unstructured condition but d = 0.08 in the structured condition; the condition × candidate-gender interaction is significant with a moderate effect size. Because assignment was random and materials identical, the reduction in gender gap is attributable to the rubric. The stated limitation: video evaluations by managers in a course setting may not capture live-interview dynamics — a field RCT is proposed.
Common pitfalls
Confusing random sampling with random assignment — only the latter creates equivalent groups
No manipulation check, so a null result cannot distinguish “no effect” from “manipulation never landed”
Underpowered designs — 20 per cell cannot reliably detect the small-to-moderate effects typical in management research
Confounded manipulations that change several things at once (longer AND interactive training vs shorter AND passive)
Differential attrition quietly destroying the equivalence randomization created
Testing many outcomes and reporting only the significant ones
Generalizing a one-hour scenario study directly to organizational policy without field evidence
What to report
Design summary: conditions, factors, between/within structure, and the randomization procedure (who generated it, how concealed)
Power analysis and the resulting target n; final n per condition; attrition by condition
Exact experimental materials and procedure (or a repository link), including manipulation and attention checks
Baseline equivalence of conditions on key characteristics
Effect sizes with confidence intervals for all pre-specified outcomes, plus assumption checks
Deviations from the pre-registered plan, and all conditions/outcomes run — not just the flattering ones
Ethics approvals, consent and debriefing procedures