empirio.ai

Data Collection: Methods, Process incl. Example

Surveys, interviews, observation or an existing dataset: which of the seven approaches suits your study, how to plan the fieldwork and what markers expect to read.

Author at empirio.ai - Maria Malzewby Maria MalzewUpdated 12 September 2026Reading time 16 min

Your research question is settled, your hypotheses are written, and then your supervisor asks the one question that decides everything: how are you going to collect your data? Get that wrong and you sit down at the end to analyse numbers that answer a different question.

Data collection is the step where you translate your research question into measurable characteristics, choose a method such as a survey, an observation, a content analysis or an experiment, and then gather the data under fixed rules written down in advance, so that the results can be analysed later and checked by other people. We show you which method suits which question, how to plan the fieldwork and what belongs in your methodology chapter.


📌 The key points

  • Data collection means gathering data systematically, not analysing it.
  • Your research question picks the method, not your timetable.
  • For 95 percent confidence and 5 percentage points, 385 responses suffice.
  • The UK Data Service gives non-PhD students End User Licence data.
  • Record fieldwork dates and response rates while the survey runs.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

Data collection: definition and how it differs from analysis

Data collection is the part of the research process in which data are gathered and measured under rules set in advance, so that they represent the object of study as accurately and with as little distortion as possible. Everything that happens to those data afterwards already counts as analysis.

Separating collection from analysis sounds like hair splitting and it is the most common reason for questions at the marking stage. Collection means gathering, analysis means calculating and interpreting. In a dissertation the two are separate chapters, and whatever you interpret early in the collection chapter is missing from the later one. How the analysis works once the data are in is covered in our article on data analysis methods in empirical research.

Measurement sits at the heart of every collection. You assign values to people, objects or texts according to rules you fixed before you started, and the rules are the part that matters. Without them you have gathered impressions rather than data.

Telling primary data from secondary data

Primary data are the data you collect yourself for your own question, with a questionnaire or an interview, for example. Secondary data come from somebody else's study, run for a different purpose: official statistics, or a dataset from a research data archive. Both routes are academically respectable, and for a dissertation on a tight timetable a careful secondary analysis often beats a survey of your own with 40 responses in it.

The UK Data Service is the archive most British students meet first. Every collection it holds sits at one of three access levels, open, safeguarded or controlled, and the level decides what you have to sign before you can download anything, as set out in its access conditions. Open data need no registration at all, safeguarded data need a registration and the End User Licence Agreement.

Students below PhD level should plan their work around safeguarded End User Licence data, according to the UK Data Service guidance on which categories of data are most suitable (checked on 12 September 2026). Secure Access and Safe Room datasets are not generally suitable for them, because those are meant for full research projects with demonstrable public benefit. A Special Licence is possible if you can justify why the End User Licence version does not fit, but the application can take a couple of months, which is longer than most dissertations have to spare.

Secondary data carry one catch you have to name in your methodology chapter: you had no say in how the questions were worded. If the other study operationalised the concept differently, the dataset measures something other than what you are asking about.

Data collection methods compared: seven approaches

Seven approaches turn up again and again in student work. They differ less in effort than in the kind of answer that ends up in your file at the end, and that is what your choice should hang on.

The overview shows what each approach is good for and what form the data take afterwards. The third column is the one most people underestimate: transcripts cannot be run through SPSS, and a matrix of numbers answers no question that starts with why.

MethodWhat it is good forWhat you end up with
Standardised surveyMeasuring frequencies and relationshipsMatrix of numbers, one row per person
Semi-structured interviewUnderstanding motives and meaningsTranscripts as running text
Focus groupWatching a group work out a positionTranscript with speaker changes
Structured observationRecording behaviour rather than self-reportObservation sheet with codes
Content analysisOpening up existing material systematicallyCoding frame with coded units
ExperimentTesting cause and effectMeasurements for each condition
Secondary analysisRe-analysing an existing datasetFinished dataset from an archive

Standardised and qualitative are not opposites but two settings of the same dial. Standardised means everyone gets the same question in the same order. Qualitative means the course of the conversation follows the person in front of you. Which setting fits depends on whether you want to count or to understand.

The standardised survey holds a special place in that list for student projects, mostly because it is the one approach you can realistically run on your own in a few weeks and still end up with enough cases to say something. What counts as a survey, and where it stops being the right tool, is set out in our article on the survey.

Which data collection method fits your research question

The method follows from the question, not from whatever looks manageable this week. Pick the tool first and write the question around it, and you find out at the next supervision meeting, when you are asked how question and design fit together.

A simple test helps. Read your research question out loud and listen to the question word, because it almost always gives away which approach belongs with it.

Illustration of a student weighing up which data collection method fits her research question

Matching question word to approach is not a rigid rule, but it holds for the great majority of student projects:

  • How many, how often, how strongly: a standardised survey or a secondary analysis.
  • Why, how does it feel, what does it mean: a semi-structured interview or a focus group.
  • What do people actually do: structured observation, because self-report drifts in a predictable direction.
  • What is in the texts, posts or records: content analysis.
  • Does A cause B: an experiment with at least two groups.

Two approaches still standing after that test? Then access to your field decides. An observation in a care home needs permission and usually an ethics approval, an interview needs people willing to give you 45 minutes, while an online survey reaches a lot of people at once without a single appointment. Leeds Beckett University calls research design the framework that guides a research project and separates a mono-method approach, quantitative or qualitative, from a mixed-methods approach in its research design guide. How your method sits inside that framework is covered in our article on research design.

💡 Tip

Write your research question and your chosen approach side by side on one sheet and show it to somebody on your course who does not know your topic. If that person cannot explain the connection in a single sentence, the fit is not there yet.

Operationalisation: turning a concept into a question

Operationalisation is the translation of a theoretical concept into something you can actually record. As long as “regular news consumption” is only a phrase, nobody can measure it, and two readers of your dissertation will picture two different things.

Operationalisation starts as early as the moment you are writing your hypotheses and runs all the way down to a single line in your questionnaire. Four stages, always in the same order, and every one of them belongs in your methodology chapter later.

  1. Define the concept. “Regular news consumption” means, in this dissertation, watching, listening to or reading the news on at least five days a week.
  2. Choose the characteristic. The measurable characteristic is the number of days per week with news use.
  3. Set the possible values. Values run from zero to seven days.
  4. Write the indicator. The questionnaire asks: “On how many days in a normal week do you use the news?”

Stage three decides what you are allowed to calculate later. A number of days is a metric quantity, while a set of A level grades is only ordinal, because the gap between an A and a B is not demonstrably the same as the gap between a D and an E. Which statistics each level of measurement allows is therefore settled here and not in the analysis.

What you have not operationalised, you have not measured. You have only asked about it.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

Planning and running your data collection in six steps

Data collection rarely fails because of the method and often fails because of the calendar. Between the final version of a questionnaire and the last response arriving there are usually several weeks in a student project, and that stretch cannot be compressed, because it belongs to other people.

The six steps below describe a quantitative online survey, because it is by far the most common case. For an interview or an observation the content of each step changes while the order stays exactly the same.

Step 1: Set a timetable with slack in it

Plan backwards from your submission date. Subtract the time for analysis and writing up first, then the fieldwork, then the pilot, then the writing of the questionnaire itself. What is left is your real starting date. For a piece of coursework, allow roughly three to four weeks for the whole collection, and for a dissertation aiming at a larger sample closer to six.

Step 2: Work out the sample size you need

The number of responses you need barely depends on the size of your population and almost entirely on the precision you want. That catches most people out, and it follows straight from the formula: above a population of roughly one million the required sample stops moving.

PopulationResponses needed
1,000 people278
10,000 people370
100,000 people383
1 million and above385

Values for a 95 percent confidence level, a margin of error of 5 percentage points and the least favourable case, a 50 to 50 split.

One thing stays true whatever the table says: a large sample makes your collection more precise but not automatically representative. Representative means the participants were selected at random and their composition matches the group you want to make a statement about. What that takes is set out in our article on the sample.

Step 3: Write the information sheet and the consent wording

Before the first question comes a short text saying who is running the study, what the data will be used for and how long taking part takes. If you collect personal data you need consent at that point which is freely given, informed and possible to withdraw. What the law asks of you is summarised in our article on data protection in surveys.

Ethical approval comes before all of that. The University of the West of England states that ethical approval is needed for all research by staff and students, undergraduate and postgraduate alike, wherever human participants are involved, and that approval must be obtained before any primary data collection for the project begins, in its guidance on why you need ethical approval. Your own university runs its own committee with its own deadlines, so ask about those early.

Step 4: Put the questionnaire together

Question order is not a formality. Questions about how people feel belong before questions about behaviour, because thinking about your media habits first colours the self-assessment that follows. Demographic questions sit at the end, where they no longer put anybody off. And every question measures exactly one characteristic: ask about satisfaction and willingness to recommend in a single sentence and you get an answer that belongs to neither.

Illustration of data being collected through a quantitative online survey and fed into an analysis

Step 5: Pilot the questionnaire with real people

Send the draft version to ten or twenty people who resemble your target group and ask them in so many words about clarity and length. A pilot is the cheapest place in the whole project to find a mistake, because once the survey is live nothing can be changed without making the responses already collected useless.

Step 6: Start the fieldwork and keep a record

Note from day one what you do: start date, end date, which channels you shared the link through, how many people started and how many finished. Those notes are not busywork, they are the raw material for your methodology chapter. Reconstruct them from memory afterwards and you write vaguely, which is exactly what a marker notices.

⚠️ Warning

Change no question once the survey is live, not even the wording. Responses from before and after the change refer to different questions and must not be analysed together. If you spot a mistake, stop and start again while only a handful of responses are in.

Common mistakes in data collection

Four mistakes turn up in dissertations again and again. All four are made while the instrument is being written, all four take minutes to avoid at that point, and none of them can be repaired once the analysis has started.

The common thread is an order of work that creeps in unnoticed: ask first, think later about what should happen to the answers. Work the analysis out in your head beforehand and you catch three of the four while reading your own draft.

Answer categories that overlap

A question about daily television with the categories “0 to 2 hours”, “2 to 4 hours” and “4 to 6 hours” looks tidy and is not. Anybody who watches exactly two hours finds two answers that fit and ticks one at random. The correct version reads “under 2 hours”, “2 to under 4 hours”, “4 to under 6 hours”. Categories have to exclude one another and cover every case between them.

Scales that promise more than they deliver

A five point agreement scale from “strongly disagree” to “strongly agree” is strictly speaking ordinal, because the gaps between the steps are not demonstrably equal. What that means for the arithmetic you are allowed to do is set out in our article on the ordinal scale. In practice such scales are often treated like an interval scale and means are calculated from them. Doing so is acceptable when you name and justify it in your methodology chapter, and a mistake when you do it quietly.

Asking about sex and gender identity as though they were one question

Sex and gender identity are two different variables, and the Government Statistical Service asks producers of statistics not to use the two terms interchangeably. No finalised harmonised standard exists for either topic in the UK: the GSS Harmonisation team published interim gender identity data harmonisation guidance on 11 December 2024, and that guidance was still the current advice when we checked it on 12 September 2026.

Until a standard arrives, the guidance points to the question asked in the England and Wales Census 2021, a voluntary question for people aged 16 and over: “Is the gender you identify with the same as your sex registered at birth?”, with the response options “Yes”, “No, enter gender identity” and “Prefer not to say”. For a student project the lesson is short. Ask only what your research question genuinely needs, keep sex and gender identity apart if you need both, and always leave a “Prefer not to say” option, exactly as the Census question does.

Leaving the choice of analysis until the data are in

Which statistical procedure you mean to use belongs before the collection, not after it. A measure of association between two variables needs different data from a comparison of means between two groups, and whether your dataset supports either is decided by the questionnaire. Our article on descriptive statistics and inferential statistics shows which procedure needs which kind of data.

How to describe data collection in your dissertation

Your methodology chapter does not explain how data collection works in general. Only what you did belongs in it, and the standard is simple: a stranger should be able to repeat your collection from your text alone.

One piece of British housekeeping before the detail. In the UK a dissertation is the piece you write for an undergraduate or taught master's degree and a thesis is the doctoral one, which is the opposite of American usage, so follow the word your course handbook uses.

A proper methods profile is easy enough to copy from a professional study. For the July 2026 round of its Opinions and Lifestyle Survey the Office for National Statistics names the collecting organisation (the ONS), the mode (an online self-completion questionnaire, with a telephone interview on request), the fieldwork dates (1 to 26 July 2026), the issued sample (8,730 households), the responding sample (3,490 individuals, a response rate of 40 percent) and the sampling approach (a random selection from people who had previously taken part in the Transformed Labour Force Survey or the Opinions and Lifestyle Survey). The full profile sits in the bulletin Public opinions and social trends, Great Britain: July 2026, released on 14 August 2026.

Illustration of a student working out where the data collection chapter belongs in her dissertation

Where the chapter sits is largely settled across British courses. Data collection follows the research design and comes before the analysis:

  1. Introduction with the research question
  2. Theory and literature review
  3. Research design with hypotheses, choice of method and sample
  4. Data collection: instrument, fieldwork, response
  5. Data analysis
  6. Findings
  7. Discussion and conclusion

Response also covers what did not work. Drop-outs, incomplete cases and a skewed age distribution are not blemishes but part of the result, and a dissertation that names them openly reads more confidently than one that keeps quiet about them. How collection fits into the sequence as a whole is covered in our article on the empirical research process.

Our reading tip: Clark, Tom, Foster, Liam, Sloan, Luke, Brookfield, Charlotte and Bryman, Alan (2026). Bryman's Social Research Methods. 7th edition. Oxford: Oxford University Press. A standard text on British social science courses and a sensible first stop for sampling and questionnaire design.

Conclusion

The best data collection is the one you finish properly in the time you have. A small, well documented survey with 120 responses carries a dissertation further than an ambitious mixed-methods design whose data are still missing three weeks before submission. Decide early, keep notes as you go, and name the limits of your collection yourself before a marker does it for you.

Where to go next


Your design is settled and all you need now are the responses?

With empirio.ai, an online survey tool from Germany, you create your questionnaire free of charge, share it with a link and see the responses as a ready analysis straight away. The data sit on servers in the EU.

Create a survey free of charge

Frequently asked questions

Data collection is the part of the research process in which data are gathered and measured under rules set in advance, so that they represent the object of study as accurately and with as little distortion as possible. Analysing those data is a separate, later step.

Seven approaches turn up regularly in student work: the standardised survey, the semi-structured interview, the focus group, structured observation, content analysis, the experiment and secondary analysis of an existing dataset. Your research question decides which one fits.

At a 95 percent confidence level and a margin of error of plus or minus 5 percentage points, 385 usable responses are enough once the population reaches one million or more. For 10,000 people the figure is 370 and for 1,000 people it is 278. A large sample makes a survey more precise but not automatically representative.

Data collection covers the systematic gathering of the data, so the instrument, the fieldwork and the response. Data analysis covers everything that happens to those data afterwards, so preparation, calculation and interpretation. In a dissertation the two are separate chapters.

Operationalisation is the translation of a theoretical concept into something measurable. Four stages lead there: define the concept, choose a measurable characteristic, set the possible values and write the question or indicator that follows from them.

For a piece of coursework, roughly three to four weeks for the whole collection is realistic, and for a dissertation with a larger target sample closer to six. Plan backwards from the submission date, because fieldwork and response depend on other people and cannot be shortened.

Five facts belong in the methodology chapter: the instrument you used, the mode of collection, the fieldwork dates with a start and an end, the sample size and the sampling approach. Add drop-outs and incomplete cases, because those are part of the result too.

You might also be interested in

Empirical Research

Sample: How to Choose Your Method and Sample Size

How many responses are actually enough? Find out which sampling methods exist, how the sample size formula works on a real example and where it stops telling you anything.