Your research question is set, your hypotheses are written, and then your advisor asks the question that decides everything: how are you going to collect your data? Answer that badly and you reach the results chapter with numbers that never touch the question you asked.
Data collection is the step where you translate your research question into measurable characteristics, pick a method such as a survey, an observation, a content analysis or an experiment, and then gather the data by fixed rules you wrote down in advance, so that the results can be analyzed later and checked by someone else. Which method fits which question, how to plan the field period, and what belongs in your methods chapter: all of it is below.
📌 Key takeaways
- Data collection means gathering data by plan, not analyzing it.
- Your research question picks the method, not your calendar.
- The USC research design guide lists 18 distinct design types.
- 385 responses cover 95 percent confidence and 5 points of error.
- ICPSR charges non-members about 825 dollars per dataset.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
Data Collection: Definition and Where Analysis Starts
Data collection is the part of a research process in which data is gathered and measured by rules set in advance, so that the result represents the thing you are studying as precisely and with as little distortion as possible. Everything you do with that data afterward already counts as analysis.
Separating collection from analysis sounds like hair splitting, and it is the most common reason for a comment in the margin. Collecting means gathering. Analyzing means calculating and interpreting. In your paper those are two separate chapters, and whatever you interpret inside the collection chapter will be missing later on. How the analysis works once the data is in is covered in our guide to data analysis methods.
Measurement sits at the heart of every collection. Measuring means assigning numbers to observations according to rules you fixed before you started, and the decisive part of that sentence is the rules. Without them you have gathered impressions, not data. The wider frame around the whole process is laid out in our overview of empirical research.
Primary Data and Secondary Data
Primary data is data you collect yourself for your own question, with a questionnaire or an interview. Secondary data comes from someone else's study, run for a different purpose: federal statistics, for example, or a dataset from a research data archive. Both routes are scientifically respectable, and for a bachelor's paper on a tight deadline a clean secondary analysis often beats your own survey with 40 responses.
Access to those archives is not automatically free. The Inter-university Consortium for Political and Social Research (ICPSR) at the University of Michigan opens part of its holdings to anyone, while its members-only datasets stay behind an institutional membership. For researchers at non-member institutions, ICPSR states that it typically charges an administration fee of approximately 825 dollars per dataset (as of September 2026, ICPSR non-member data access policy). Ask your university library first, because most large U.S. universities are already members.
Secondary data carries one catch you have to name in your methods chapter: the wording of the questions was not yours. When the original study operationalized a concept differently than you do, the dataset measures something other than what you are asking about.
Seven Data Collection Methods Compared
Seven methods show up again and again in student research. They differ less in effort than in the kind of answer that ends up in your file, and that is what your choice should hang on.
The overview below shows what each method is good for and what form the data takes afterward. The third column is the one most people underestimate: transcripts cannot be run through SPSS, and a matrix of numbers answers no question that begins with why.
| Method | What it is good for | What you end up with |
|---|---|---|
| Standardized survey | measuring frequencies and relationships | numeric matrix, one row per person |
| Semi-structured interview | understanding motives and meanings | transcripts as running text |
| Focus group | watching a group negotiate meaning | transcript with speaker turns |
| Structured observation | recording behavior instead of self-report | observation sheet with codes |
| Content analysis | opening up existing material systematically | coding scheme with coded units |
| Experiment | testing cause and effect | measurements per condition |
| Secondary analysis | re-analyzing an existing dataset | finished dataset from an archive |
Standardized and qualitative are not opposites. Standardized means everyone gets the same question in the same order. Qualitative means the course of the conversation follows the person in front of you. Which setting matches your interest depends on whether you want to count or to understand.
Among the seven, the survey is the one we see most often in student projects, which is why it gets an article of its own. What a survey is and where it runs out of road is covered in our article on the survey.
Which Data Collection Method Fits Your Research Question
The method follows from the question, not from what happens to be feasible this month. Pick the tool first and write the question around it, and the mismatch surfaces the moment your advisor asks how design and question fit together.
One quick test helps: read your research question out loud and listen to the question word. Almost every time, the question word tells you which method belongs to it.
Research Designs: 18 Types, Not Three
Research design is a wider field than most course readers let on. The USC Libraries guide to research designs, part of the university's writing guide for the social sciences, describes 18 separate design types, from action research and case study through cohort, longitudinal and photovoice designs (USC Libraries, Types of Research Designs). A three-way split into qualitative, quantitative and mixed methods is a starting point, not the map.

Inside quantitative work the real axis runs between descriptive and experimental. USC puts it plainly: a descriptive design answers who, what, when, where and how, and cannot conclusively establish why. An experimental design specifies an experimental group and a control group, gives the independent variable to one and not the other, and measures both on the same dependent variable. Nothing short of that structure supports a causal claim.
Match the Question Word to the Method
The mapping from question word to method is not a rigid rule, and it holds for most student projects. Read your question once more and see which of the five lines below it lands on, because the answer usually settles the choice in a few seconds.
- How many, how often, how strongly: standardized survey or secondary analysis.
- Why, how do people experience, what does it mean: semi-structured interview or focus group.
- What do people actually do: structured observation, because self-report drifts in a predictable direction.
- What is in the texts, posts or records: content analysis.
- Does A cause B: experiment with at least two groups.
When two methods survive that test, access to your field decides. An observation in a nursing home needs permission and usually an IRB review, and a 45-minute interview needs people willing to give you 45 minutes. An online survey reaches many people at once without a single appointment. How method and overall setup hang together is explained in our guide to research design.
💡 Tip
Write your research question and your chosen method side by side on one sheet and show it to a classmate who does not know your topic. When that person cannot explain the connection in one sentence, the fit is not there yet.
Operationalization: Turning a Concept Into a Question
Operationalization is the translation of a theoretical concept into something you can actually record. As long as “regular news consumption” is only a phrase, nobody can measure it, and two readers of your paper will picture two different things.
Operationalization starts as early as the moment you are writing your hypotheses and runs all the way down to the single line in your questionnaire. Four steps, every time, and each one belongs in your methods chapter later.
- Define the concept. “Regular news consumption” means, in this paper, watching, listening to or reading news on at least five days per week.
- Name the characteristic. The measurable characteristic is the number of days per week with news use.
- Set the possible values. The range runs from zero to seven days.
- Write the indicator. The questionnaire item reads: on how many days in a normal week do you use news?
Step three decides what you are allowed to calculate later. A count of days is a metric quantity. School grades are only ordinal, because the distance between an A and a B is not demonstrably the distance between a D and an F. The level of measurement you settle on here, and not the one you wish for in the analysis chapter, sets the statistics you may use.
What you have not operationalized, you have not measured. You have asked about it.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
Planning and Running a Data Collection: Six Steps
A data collection rarely fails on method and often fails on the calendar. Between the finished questionnaire and the last response arriving, student projects usually see several weeks pass, and those weeks cannot be compressed, because other people control them.
The six steps below describe a quantitative online survey, since that is the most common case. For an interview or an observation the content of the steps changes, while the order stays the same.
Step 1: Set a Schedule With Slack in It
Plan backward from your due date. Subtract the time for analysis and writing first, then the field period, then the pretest, then the time it takes to create the questionnaire. What is left over is your real start date. For a seminar paper, count on roughly three to four weeks for the whole collection; for a master's thesis with a representative target size, closer to six.
Step 2: Calculate Your Sample Size
The number of responses you need barely depends on the size of your population and almost entirely on the precision you want. A lot of people find that counterintuitive, and it follows straight from the formula: past a population of roughly one million, the required sample stops moving.
| Population | Responses needed |
|---|---|
| 1,000 people | 278 |
| 10,000 people | 370 |
| 100,000 people | 383 |
| 1 million and above | 385 |
Values for a 95 percent confidence level, a margin of error of plus or minus 5 percentage points, and the worst case of a 50/50 split.
One point stays true whatever the size: a large sample makes your collection more precise, not automatically representative. Representative only happens when participants were selected at random and their composition matches the group you want to make a statement about. What that takes is laid out in our article on the sample.
Step 3: Write the Cover Text, the Consent and the IRB Question
Before the first question comes a short text saying who is running the study, what the data will be used for, and how long participation takes. When you collect personal information, you need consent at that point: voluntary, informed and revocable. What the rules require of you is summarized in our guide to data privacy in surveys.
Whether your project needs IRB review turns on two federal definitions, not on how sensitive your topic feels. Under 45 CFR 46.102(e)(1), a human subject is a living individual about whom an investigator obtains information through intervention or interaction and then analyzes it, or about whom the investigator obtains, uses or generates identifiable private information (eCFR, 45 CFR 46.102). Harvard's Committee on the Use of Human Subjects states the test in one line: when a project meets both the definition of research and the definition of human subjects, it counts as regulated research and IRB review is needed (Harvard CUHS, Do You Need IRB Review). Ask your own IRB early, because the determination is theirs to make.
Step 4: Create the Questionnaire
Question order is not a formality. Items about how people feel belong ahead of items about what they do, because thinking about media habits first colors the self-assessment. Demographic questions go at the end, where they no longer scare anyone off. And every question measures exactly one thing: ask about satisfaction and willingness to recommend in one sentence, and the answer belongs to neither.

Step 5: Run a Pretest With Real People
Send the draft version to ten or twenty people who resemble your target group and ask them explicitly about clarity and length. A pretest is the cheapest bug hunt in the whole project, because once the survey is live, nothing can be changed without making the responses you already have unusable.
Step 6: Open the Field and Document It
Write down what you do from day one: start date, end date, which channels you shared through, how many people started and how many finished. Those notes are not busywork. They are the raw material for your methods chapter, and anyone who reconstructs them afterward writes vaguely, which is exactly what a grader notices.
⚠️ Caution
Once the survey is live, change nothing, not even a single word of a question. Responses from before and after the change belong to different questions and must not be analyzed together. When you spot an error and only a few responses are in, stop the collection and start over.
Common Mistakes in Data Collection
Four mistakes turn up in student papers over and over. All four are made while creating the instrument, all four take a couple of minutes to avoid at that stage, and none of them can be repaired during analysis.
The common thread is an order of operations that creeps in unnoticed: ask first, work out later what should happen to the answers. Run the analysis in your head before you collect, and three of the four show up while you are still reading your own draft.
Answer Categories That Overlap
A question about daily TV time offering “0 to 2 hours”, “2 to 4 hours” and “4 to 6 hours” looks tidy and is not. Somebody who watches exactly two hours finds two answers that fit and checks one at random. The correct wording is “under 2 hours”, “2 to under 4 hours”, “4 to under 6 hours”. Categories have to exclude each other and cover every case between them.
Scales That Promise More Than They Deliver
A five-point agreement scale running from “strongly disagree” to “strongly agree” is strictly speaking ordinal, because the distances between the steps are not demonstrably equal. What that means for the arithmetic you are allowed to do is set out in our article on the ordinal scale. In practice such scales are often treated as interval and means are calculated. Doing so is defensible when you name and justify it in your methods chapter, and a mistake when you do it quietly.
Demographic Categories That Force a Wrong Answer
Federal standards for demographic questions changed in 2024, and plenty of questionnaire templates still carry the old ones. On March 29, 2024, the Office of Management and Budget published its revised Statistical Policy Directive No. 15, which supersedes the 1997 race and ethnicity standards. The revision requires a single combined race and ethnicity question that allows multiple responses, and it adds Middle Eastern or North African as a minimum reporting category, separate and distinct from White (Federal Register, 89 FR 22182). Existing federal collections have until March 28, 2029 to fall in line.
Your seminar paper is not a federal statistical program, and the lesson still applies directly. A demographic question with too few categories forces people into an answer that is wrong for them, and every one of those answers pollutes the variable you wanted to analyze. Give respondents categories that actually describe them, allow more than one answer where the concept allows it, and offer a “prefer not to answer” option so that silence is a choice rather than a guess.
Looking for an Analysis Method After the Fact
Which statistical procedure you plan to use belongs before the collection, not after it. A measure of association between two variables needs different data than a comparison of means between two groups, and whether your dataset supports either is decided by the questionnaire. Our article on descriptive statistics and inferential statistics shows which procedure assumes which kind of data.
Describing Your Data Collection in a Thesis or Dissertation
Your methods chapter reports only what you did, not how data collection works in general. The standard is simple: a stranger should be able to repeat your collection from your text alone, without asking you a single follow-up question.
A good model for that kind of profile is a professional survey organization. Pew Research Center publishes a full methodology statement with every report, and the statement for its May 11, 2026 report on the problems Americans see facing the nation names all five pieces you need.
Pew gives the instrument (Wave 192 of the American Trends Panel), the field dates (April 20 to 26, 2026), the sample (5,103 respondents out of 5,898 sampled, an 87 percent survey-level response rate), the modes (online with n=4,900 and live telephone with n=203, fielded by SSRS) and the sampling approach (address-based sampling from the U.S. Postal Service's Computerized Delivery Sequence File, with the adult in the household who has the next birthday selected). The margin of sampling error is reported as plus or minus 1.6 percentage points (as of September 2026, Pew Research Center methodology, Wave 192). Your methods chapter needs the same five items, with your numbers.
Where the Data Collection Chapter Sits in the Paper
Where the chapter sits in the paper is largely settled. Data collection follows the research design and comes before the analysis:
- Introduction with the research question
- Theory and state of research
- Research design with hypotheses, choice of method and sample
- Data collection: instrument, field period, response
- Data analysis
- Results
- Discussion and conclusion

Response also covers what did not work. Break-offs, incomplete cases and a lopsided age distribution are not blemishes, they are part of the result, and a paper that names them openly reads more confident than one that hides them. Where collection sits in the whole sequence is described in our guide to the empirical research process.
Our reading tip: the USC Libraries guide Organizing Your Social Sciences Research Paper covers all 18 design types and adds a separate page on design flaws to avoid, which is the closest thing to a checklist for a methods chapter you will find for free.
Conclusion
The best data collection is the one you finish cleanly in the time you have. A small, well-documented survey with 120 responses carries a thesis further than an ambitious mixed-methods design whose data never arrives. Decide early, take notes as you go, and name the limits of your collection yourself before a grader does.
What to Read Next
- Want to know which method fits your question? Qualitative and quantitative research methods
- Putting your questionnaire together right now? How to create a questionnaire
- Looking for people to take your survey? Tips for online surveys
- Writing an empirical thesis? Empirical research in final papers
Your collection is planned. Now you just need the responses.
With empirio.ai, an online survey tool from Germany, you create your questionnaire for free, share it with a link and watch the answers arrive as a live analysis. The data stays on servers in the EU.
