Surveys, interviews, focus groups, observation, experiments and archives — what each is good for, what it costs, and how to keep bias out of the pipeline.
Methodology step
Matching method to question
Data collection is where design meets reality. The method must fit the question (attitudes vs behaviour vs meaning), the population (literacy, access, sensitivity) and the budget. Strong studies increasingly combine sources — a survey for breadth plus records for objective outcomes — because every single method has a characteristic weakness that another can cover.
Surveys / questionnaires
Structured self-report at scale; cheap per respondent; ideal for attitudes, perceptions and reported behaviour. Weaknesses: response biases (social desirability, acquiescence), nonresponse, and the gap between reported and actual behaviour.
Interviews
Depth, probing and flexibility (structured → semi-structured → unstructured); handles complexity and sensitive topics with rapport. Costly per participant; interviewer effects; needs skilled interviewing and verbatim transcription.
Focus groups
6–10 participants in moderated discussion; efficient for the range of views, shared norms and natural language; the interaction itself is data. Risks: dominant voices, conformity pressure, unsuitable for confidential topics.
Observation
Watch behaviour directly (participant or non-participant; structured checklists or open field notes). Captures what people do rather than say. Costs: time, access, observer effects (behaviour changes under watch), and inference limits about motives.
Experimental tasks & measures
Outcomes recorded under controlled conditions — choices, performance, response times. High internal validity; artificiality and demand effects are the standing worries.
Secondary / archival data
Existing records: HRIS, sales, financial filings, government statistics, review platforms. Cheap, longitudinal, non-reactive — but collected for other purposes, with definitions, gaps and access constraints you inherit.
Designing the survey encounter
Pilot everything: comprehension, completion time, device rendering; revise before launch, not after
Question order matters — general before specific; sensitive and demographic items last
Avoid double-barrelled, leading and hypothetical-heavy items; every item should map to a construct in your model
Use established response formats consistently; label all scale points
Track and report response rate, and compare respondents to the population (or early vs late responders) for nonresponse bias
Common-method bias — the standing threat in survey research
When the same respondents report both predictors and outcomes, in the same questionnaire, at the same moment, part of any observed correlation can come from the shared method rather than the constructs: consistency motives, mood, item wording similarity, social desirability. This common-method variance can inflate (or occasionally deflate) relationships and is a routine examiner objection to single-source cross-sectional surveys.
Procedural remedies (best, because they prevent rather than diagnose): obtain predictor and outcome from different sources (employee self-report + supervisor rating + system data); separate measurements in time; guarantee anonymity to reduce social desirability; separate scale blocks and vary response formats; keep items concrete and unambiguous
Statistical diagnostics (weaker, but expected): Harman's single-factor test is common but insensitive — do not rely on it alone; report an unmeasured latent method factor or a theoretically unrelated marker variable where feasible
Honest reporting: state which remedies were used and calibrate conclusions — no statistical patch fully rescues a design where every construct came from one questionnaire in one sitting
Working with secondary and archival data
Interrogate provenance: who collected it, why, with what definitions and coverage — misaligned definitions are the classic trap
Check consistency over time: systems, categories and thresholds change and mimic real trends (instrumentation)
Document cleaning and matching decisions; keep raw and processed versions separate
Confirm permission and privacy compliance before extraction, not after
Ethics runs through everything
Informed consent that names recording, storage and use; extra care with employer-mediated recruitment where refusal may feel unsafe
Anonymity vs confidentiality — promise only what your design can deliver (linked multi-wave data is not anonymous)
Secure storage and de-identification; aggregate reporting where individuals could otherwise be recognized
Institutional ethics approval obtained before the first participant is contacted
Common pitfalls
Instrument finalized without piloting — ambiguities discovered in the data are permanent
All constructs measured from one source at one time with no remedy, then causal-sounding conclusions drawn
Focus groups used for topics people will not discuss in front of peers
Observation without a protocol, producing anecdotes rather than data
Archival variables accepted at face value without checking definitional drift
Consent and approval treated as paperwork after the design is frozen