Evergreen guide

Data collection methods

Ten methods, compared on what they cost, how long they take, how many people you need and the bias each one carries. Then the parts most guides skip: how to size a sample, how to write questions that don't decide the answer for you, and which method to pick when you have three weeks and no budget.

Every section links to a form you can copy and use — 10 research templates built for this guide, from a 625-template library. Last updated .

What data collection actually means

Data collection is the deliberate, documented capture of evidence against a question you wrote down first. The order matters: a question, then a method chosen because it can answer that question, then an instrument — a questionnaire, an interview guide, an observation sheet — and only then respondents. Most weak studies invert it. They start with a tool somebody already had, collect whatever it produces, and look for a question the data happens to fit.

Three decisions determine whether the result is usable. The unit of analysis: are you describing people, sessions, transactions or days? Mixing them is the most common reason two analysts get different numbers from the same file. The instrument: fixed wording, fixed options and fixed units, frozen before fieldwork opens. The selection rule: how people or records got into the dataset, which decides what you are entitled to claim at the end.

Everything after that is execution. A clean instrument with a broken selection rule produces confident nonsense; a rough instrument with an honest selection rule produces a useful caveat. Both are better than a survey that was never written down, sent to whoever answered, and then reported as a percentage.

Primary vs secondary data

Primary data is collected by you, for your question, with an instrument you control. You choose the wording, the sample and the timing, so the data fits the question exactly — and you pay for all of it in time, recruitment and fieldwork.

Secondary data was collected by somebody else for their own purposes: official statistics, published research, industry reports, administrative records, platform analytics. It is faster, cheaper, usually far larger and frequently better collected than anything you could field alone. The catch is definitional. Their age bands, reporting periods and thresholds were chosen for their question, not yours, and a mismatch you notice late can invalidate an entire analysis.

Primary and secondary data compared
DimensionPrimarySecondary
Fit to your questionExact — you wrote the questionsApproximate — their definitions
CostYour time, recruitment, incentivesUsually free or licensed
SpeedWeeks to monthsHours to days
ScaleLimited by budgetOften whole populations
Control of qualityYours to get right or wrongFixed; you can only assess it
Main riskBias you introducedDefinition mismatch and revision

In practice most good projects use both: secondary data to establish context and benchmarks, primary data for the specific thing nobody has measured. Keep a Secondary Data Log from the first download — publisher, reference period, licence, access date — because the provenance question always arrives later than the download.

Quantitative vs qualitative methods

Quantitative methods produce values you can count, compare and test: how many, how often, how much, how likely. Qualitative methods produce language and observation you can interpret: how something works, why somebody did it, what they thought was happening. The distinction is not rigour — both can be done well or badly — it is the kind of claim each one supports.

The failure modes are mirror images. Quantitative work fails by measuring a proxy precisely and calling it the thing itself. Qualitative work fails by generalising twelve conversations into a population claim. Mixed-method designs exist to cover for each other: interviews first to find out what the questions should be, a survey to find out how common each answer is, then a handful of interviews again to explain the result that surprised you.

Use quantitative when…

  • You already know the possible answers and need their frequencies.
  • You need to compare groups, sites or time periods.
  • A decision hinges on a threshold or an effect size.
  • Someone will ask for a confidence interval.

Start with the Research Survey Template or the Likert Scale Survey.

Use qualitative when…

  • You don't yet know what the options should be.
  • The mechanism matters more than the frequency.
  • People's own words are the evidence.
  • The behaviour only makes sense in its setting.

Start with the Interview Guide Form or the Observation Log Form.

The ten methods, compared

Cost and speed below assume you run the fieldwork yourself with no agency. "Main bias" is the failure this method is most exposed to — not the only one, but the one that most often survives into a published finding.

Ten data collection methods compared on cost, speed, sample, bias and template
MethodBest forCostSpeedTypical sampleMain biasStart with
Surveys & questionnairesquantitativeMeasuring how common something is across a populationLowFast100–1,000+Non-response bias — the people who answer differ from those who don'tResearch Survey Template
InterviewsqualitativeUnderstanding why people behave the way they doMediumSlow8–30Interviewer effect — people answer the person, not the questionInterview Guide Form
Focus groupsqualitativeSurfacing shared language, norms and disagreementHighModerate3–6 groups of 6–8Groupthink — one confident voice sets the roomFocus Group Screener
ObservationmixedBehaviour people cannot accurately report about themselvesMediumSlow10–60 windowsObserver effect — being watched changes the behaviourObservation Log Form
Diary studies & experience samplingmixedHow something changes over time, in contextMediumSlow10–40 people × 7–14 daysAttrition — compliance drops as the study runsDiary Study Check-in
Experiments & A/B testsquantitativeEstablishing that a change caused an effectMediumModerateHundreds to thousands per armNovelty effects and peeking at results before the test endsPre & Post Test Form
Measurement & instrument readingsquantitativeObjective quantities that don't depend on opinionMediumModerateAs many readings as the protocol demandsInstrument drift and inconsistent units between recordersData Collection Sheet
Tests & assessmentsquantitativeShowing learning or capability, especially before and afterLowFast20–500Practice effects when the same test is reused too soonPre & Post Test Form
Administrative & operational recordsquantitativeLong time series you never have to fieldLowFastEverything you haveDefinitions built for operations, not for your research questionData Collection Sheet
Secondary data & desk researchmixedContext, benchmarks and questions too big to field yourselfLowFastWhole populationsDefinition mismatch between their categories and yoursSecondary Data Log

Surveys & questionnaires

The same fixed questions put to many people so answers can be counted and compared.

Watch for: non-response bias — the people who answer differ from those who don't.

Use the Research Survey Template

Interviews

A guided one-to-one conversation that follows the participant's reasoning.

Watch for: interviewer effect — people answer the person, not the question.

Use the Interview Guide Form

Focus groups

A moderated group discussion where participants react to each other.

Watch for: groupthink — one confident voice sets the room.

Use the Focus Group Screener

Observation

Recording what people actually do, in the setting where they do it.

Watch for: observer effect — being watched changes the behaviour.

Use the Observation Log Form

Diary studies & experience sampling

Short repeated entries collected from the same people over days or weeks.

Watch for: attrition — compliance drops as the study runs.

Use the Diary Study Check-in

Experiments & A/B tests

Two or more conditions assigned deliberately so the difference can be attributed.

Watch for: novelty effects and peeking at results before the test ends.

Use the Pre & Post Test Form

Measurement & instrument readings

Physical readings taken with a calibrated instrument to a fixed protocol.

Watch for: instrument drift and inconsistent units between recorders.

Use the Data Collection Sheet

Tests & assessments

A scored instrument that measures knowledge or ability rather than opinion.

Watch for: practice effects when the same test is reused too soon.

Use the Pre & Post Test Form

Administrative & operational records

Data your own systems already produce as a by-product of doing the work.

Watch for: definitions built for operations, not for your research question.

Use the Data Collection Sheet

Secondary data & desk research

Data somebody else collected, published and licensed for reuse.

Watch for: definition mismatch between their categories and yours.

Use the Secondary Data Log

Sampling methods

Sampling is the selection rule, and it decides what your numbers are allowed to mean. Probability sampling gives every member of the population a known, non-zero chance of selection, which is what makes generalisation and margins of error legitimate. Non-probability sampling does not — it can still be the right choice, but the claim at the end has to shrink to match.

Probability

Simple random

Every member of the population has an equal chance of selection, drawn at random from a full list.

Use when: You have a complete, accurate list of the population.

Weakness: Small subgroups can be missed entirely by chance.

Probability

Systematic

Pick every nth record from an ordered list after a random start.

Use when: The list is long and has no hidden periodic pattern.

Weakness: A repeating pattern in the list order maps straight into your sample.

Probability

Stratified

Split the population into strata (age, region, plan) and sample randomly inside each.

Use when: You need reliable estimates for subgroups, not just the whole.

Weakness: You need to know the strata sizes in advance.

Probability

Cluster

Sample whole naturally occurring groups — schools, wards, branches — then everyone (or a sample) inside them.

Use when: The population is geographically spread and travel costs dominate.

Weakness: People inside a cluster resemble each other, so effective sample size shrinks.

Non-probability

Convenience

Ask whoever is available and willing.

Use when: Piloting an instrument, or exploratory work with no inference claim.

Weakness: No basis for generalising to a population — margin of error is not meaningful.

Non-probability

Purposive

Deliberately select people who have the experience you need to understand.

Use when: Qualitative work where depth matters more than spread.

Weakness: The selection reflects the researcher's assumptions.

Non-probability

Quota

Fill preset counts per group until each quota is complete.

Use when: Commercial research needing a demographically balanced group fast.

Weakness: Selection inside each quota is still non-random.

Non-probability

Snowball

Ask participants to refer others like them.

Use when: Hard-to-reach or hidden populations with no usable list.

Weakness: Samples travel along social ties, so isolated cases never appear.

For quota and purposive designs, the selection happens in the screener. The Focus Group Screener is built for exactly this: it asks the qualifying questions before it explains the study, so nobody can work out which answer gets them in.

How big does your sample need to be?

For a quantitative estimate, sample size comes from three inputs: how precise you need to be (margin of error), how sure you need to be that the true value sits inside that range (confidence level), and how large the population is. The standard result surprises people — for any large population, ±5% at 95% confidence needs about 385 completed responses whether the population is fifty thousand or fifty million. Population size only matters when it is small.

Sample size calculator

Enter your population and the precision you need. The result is the number of completed responses required, and how many invitations that takes at your expected response rate.

Leave blank if the population is effectively unlimited.

Completed responses needed371Invitations to send1,484Ignoring population size385

Assumes a probability sample and the most conservative expected proportion (50%). A self-selected sample has no valid margin of error, however many responses it collects.

Qualitative work sizes differently. There is no formula; you sample until new sessions stop producing new themes. In practice that is usually 8–15 interviews for a fairly homogeneous group, more if you are comparing distinct populations, and 3–6 focus groups before the discussion starts repeating itself. For diary studies, plan for attrition rather than pretending it won't happen: recruit 10–40 people, expect a fifth to drift, and check compliance on day three rather than day ten.

Writing questions that don't skew the answer

Most bias in survey data is introduced by the person who wrote the questionnaire, not by the people who answered it. Six rules prevent the majority of it.

Question design rules with a poor and a better example
RuleWeakBetter
One idea per questionWas the service fast and friendly?How fast was the service? / How friendly was the service?
No leading wordingHow much did you enjoy our award-winning support?How would you rate the support you received?
Balanced scalesGood / Very good / ExcellentStrongly disagree / Disagree / Neither / Agree / Strongly agree
Answerable time framesHow often did you use it last year?How many times did you use it in the last 7 days?
Exhaustive, exclusive options0–10, 10–20, 20–300–9, 10–19, 20–29, 30 or more
An honest way outIncome (required)Income — including a 'prefer not to say' option

Likert scales, done properly

A Likert item is a statement plus a symmetric agreement scale — not any old rating question. Keep two negative points, a genuine neutral midpoint and two positive points, use identical wording on every item, and reverse-word a couple of statements in each block so that respondents ticking straight down the page become visible in the export. Offer "not applicable" separately from the midpoint: "no opinion" and "does not apply to me" are different answers. Five points is enough for most general samples; seven adds variance for expert respondents and little else. When you report, publish the distribution as well as any mean — a bimodal set of answers has no meaningful average, and 3.72 on a five-point scale implies precision the instrument does not have. The Likert Scale Survey ships with all of this built in, reverse-worded items included.

Which method should you use?

Method choice is mostly constraint arithmetic: what you need to know, who you can reach, what you can spend and how long you have. Answer four questions and the chooser names a design plus the claim it will not support.

Which method should you use?

Four questions. The answer names a primary method, a method to pair it with, and the claim the design will not support.

Recommended design

Surveys & questionnaires + Secondary data & desk research

A structured questionnaire on a probability sample gives you countable answers, and published data tells you whether your sample looks like the population.

Honest limit: A margin of error only means something for a probability sample. A self-selected sample has no valid margin of error.

Tools and instruments

You need four capabilities, and they are less glamorous than the software category suggests: an instrument you can version, capture without re-typing, storage you can export from, and an audit trail of consent.

  • The instrument. A form builder for questionnaires, screeners, consent, logs and field sheets. The advantage over a document is that logic runs at capture time — a follow-up appears only when the answer needs one, and required fields cannot be skipped.
  • Capture. Anything that avoids transcription. Paper and spreadsheets both introduce copying errors; a form on a phone that queues offline and syncs later does not.
  • Analysis. A spreadsheet is enough for frequencies and cross-tabs. R, Python, SPSS or Stata for anything inferential. Qualitative coding can be done in a spreadsheet up to about thirty transcripts before dedicated software pays for itself.
  • Governance. Consent records stored apart from the data, keyed by participant code, plus a stated retention period and a route to withdraw.

HelloForms covers the first two and the last: conditional logic, offline-tolerant capture, CSV and PDF export, and a Research Consent Form with each permission recorded separately. For readings taken in the field, the Data Collection Sheet fixes units in the labels so a day of measurements cannot arrive in three different scales.

Ethics, consent and data protection

None of this is legal advice, and your ethics committee or data protection lead has the final word. But five points come up in almost every review.

  • Consent is several decisions, not one. Taking part, being recorded, being quoted, having data reused and being contacted again are separate permissions. Bundling them into a single tick box is not informed consent, and it costs you the participants who would have said yes to three out of five.
  • Collect less than you can. Every field you add is a field you must justify, store and eventually delete. If an answer will not change what you do, it does not belong on the form.
  • Anonymity and confidentiality are different promises. Anonymous means you cannot identify the respondent even if you want to. Confidential means you can but you won't. Say which one you are offering, and never say 'anonymous' on a form that carries an email field.
  • State retention and withdrawal in concrete terms. Give a period, not 'as long as necessary', and name the point after which data can no longer be pulled out of an analysis. A vague promise is worse than an honest limit.
  • Special category data raises the bar. Health, ethnicity, religion, sexuality, biometrics and political opinion carry extra obligations under UK and EU data protection law. If you are collecting them, you need a documented lawful basis before you field the form.

Nine mistakes that ruin a dataset

  • Treating a self-selected sample as representative. Report it as what it is — the people who chose to answer — and compare their characteristics with the population you claim to describe.
  • Quoting a margin of error on a convenience sample. Margins of error assume random selection. Without it, report counts and distributions and drop the ± figure.
  • Changing the instrument mid-study. Freeze the wording once fieldwork starts. If you must change it, treat the data as two datasets and say so.
  • Mixing description with interpretation in field notes. Two boxes: what was seen, and what the observer thinks it meant. Descriptions can be re-analysed; conclusions cannot.
  • Leaving units to free text. Put the unit in the field label and accept numbers only. Half your readings otherwise arrive in the wrong scale.
  • Losing the pairing between rounds. Issue matching codes at enrolment so pre and post entries can be paired without storing names.
  • Asking demographics first. Move age, income and ethnicity to the end. At the top they raise abandonment before anyone knows why you're asking.
  • Not recording the access date on a source. Log publication date and access date separately. Online sources change without notice.
  • Peeking at an experiment and stopping when it looks good. Fix the sample size and the end date before launch. Early stopping manufactures significance.

Frequently asked questions

Forms for every method on this page

Copy any of these into your workspace, change the wording, and publish. Conditional logic, required fields and exports are already set up.

Want the whole set in one place? Browse Research & Data Collection, or read the conditional logic guides to add branching of your own.