In 1936 an American magazine mailed more than ten million ballots to predict a presidential election. Over two million people filled one out, and the forecast still named the wrong winner, missing by roughly 20 percentage points.
You determine a sample in three steps: define the population, choose a sampling method, and calculate the size that follows from both. Only random sampling lets you generalize to the wider population, and the sample size formula holds for random sampling alone. By the end you will know how many participants your project needs and how to justify that choice in your methods section.
📌 The key points at a glance
- A sample is the subset of a larger population you actually study.
- Only random sampling supports conclusions about the wider population.
- At 95 percent confidence and a 5 percent margin, 385 people suffice.
- Above one million people, population size no longer changes the answer.
- The American Community Survey samples about 3.5 million addresses yearly.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeWhat is a sample?
A sample is a subset drawn from a larger group, the population, and studied on behalf of that group. Findings from the subset are generalized to the whole group, provided the subset was selected through a recognized statistical procedure rather than by convenience.
The population covers everyone or everything you want to make a statement about. Statistics writes the population as capital N and the sample as lowercase n. Studying every element of the population is a census. Studying part of it is a sample survey, and in practice that is almost always the only workable route.
Official statistics show how small a sample can be. The U.S. Census Bureau randomly selects about 3.5 million addresses each year for the American Community Survey, a fraction of all addresses in the country, and responding is required by law under Title 13 of the U.S. Code (as of August 2026). From that fraction come the figures on income, commuting and housing that guide federal funding decisions. Your own empirical research works on the same principle, on a smaller scale.
A census is the exception. A justified selection is the normal case.
Why a large sample is not automatically representative
A large sample is not automatically representative, because the number of responses says nothing about how those responses came about. If the selection method distorts the composition, every additional response only measures the wrong result more precisely.
The best known evidence comes from the magazine The Literary Digest. Before the 1936 election it mailed over ten million ballots, mostly to addresses taken from telephone directories and automobile registration lists. More than 2.3 million came back, a response rate just under a quarter. The forecast gave 55 percent to Alf Landon and 41 percent to Franklin D. Roosevelt. Roosevelt won, with 60.8 percent of the popular vote (The American Presidency Project, University of California Santa Barbara).
Almost every textbook offers the same explanation: only wealthier households owned a telephone and a car, so the selection was skewed. The case is less tidy than that. Peverill Squire analyzed a follow-up survey in Public Opinion Quarterly in 1988 and found that even people who owned both a car and a telephone voted mainly for Roosevelt. The second error sat in the response: Landon supporters returned the ballot far more often. How the total error splits between the two causes is still disputed. Squire (1988) weights the address lists more heavily, Dominic Lusinchi (2012) the response.
One uncomfortable rule follows for your own survey: who replies shapes the result as strongly as who was invited. Whether a survey ends up representative therefore depends on the method and the response, not on hitting a minimum number.
Sampling methods: the three routes to your selection
Sampling methods fall into three groups: arbitrary selection, purposive selection and random selection. Only random selection supports a statistically grounded conclusion about the population, because only there does every element have a known chance of being drawn.
Which group applies to you is rarely settled by the ideal and almost always by a practical question: do you have a complete list of your population at all? No such list exists for “everyone who bikes in Chicago”, so genuine random selection is off the table. The honest answer in your methods section is then that you selected purposively or arbitrarily.
| Method | How people are selected | When it fits |
|---|---|---|
| Arbitrary selection | whoever happens to be reachable | early exploration, pilot testing |
| Purposive selection | by criteria set in advance | extreme cases, rare groups |
| Random selection | by a random procedure from a list | conclusions about the population |
The three most common random samples
Within random selection the methods literature distinguishes three variants that differ sharply in effort. All three meet the same condition, that chance decides who takes part. They simply apply that chance at different points in the procedure, and that changes how precise the final result becomes.
- Simple random sample: every element of the population is drawn directly and with equal probability.
- Stratified random sample: the population is first split by a characteristic, then drawn within each stratum.
- Cluster sample: whole groups such as school classes are drawn, and everyone inside is surveyed.
For student projects the stratified variant is often the most sensible, because it keeps an important group from ending up too small by pure chance. How to reach enough people for the draw is covered in the guide to finding survey participants.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeHow to calculate sample size: formula, values and an example
Sample size follows from four inputs: the size of the population, the margin of error you accept, the confidence level you want, and the expected spread of answers. Together they give the minimum number of people who need to complete your survey.
The four inputs in detail:
- Population (N): everyone you want to make a statement about.
- Margin of error (e): how far your result may sit from the true value, commonly 5 percent.
- Confidence level: how reliably that interval captures the true value, commonly 95 percent.
- z-score: the confidence level expressed as a number, 1.645 at 90 percent, 1.96 at 95 percent, 2.576 at 99 percent.
- Spread (p): the expected distribution of answers, set to 0.5 without prior knowledge, the least favorable case.
Two steps turn those inputs into a number. The first gives the value for a very large population, the second corrects it downward when your group is small:
n0 = z² × p × (1 − p) / e²
n = n0 / (1 + n0 / N)
A worked example: you survey the students at your university, 12,000 people in total. At 95 percent confidence (z = 1.96), a 5 percent margin of error (e = 0.05) and p = 0.5, the first step gives 384.16 and the second 372.2. Rounded up, you need 373 completed responses. Because a share of respondents always drops out partway, send the questionnaire to considerably more people than that.
When the sample size formula does not apply
The formula assumes random selection. If you share your survey through an Instagram story, a group chat or a bulletin board in your department, chance is not deciding who takes part. Willingness among the people who already see you is. No margin of error can be calculated for a convenience sample of that kind, not even with 500 responses.
⚠️ Careful
385 responses from an Instagram story say nothing about the United States, only something about your reach. Do not put a margin of error in your paper in cases like that. Name the method instead, limit the claim to the group you actually reached, and explain why that is enough for your research question.
How many participants do you really need?
For most projects the sample you need sits between 80 and 400 people. At 95 percent confidence and a 5 percent margin of error it is 385 people once the population passes one million, and considerably fewer as soon as your group is manageable.
The surprising part is how little population size matters. Between one million and 340 million people the required sample does not change at all. Anyone who inflates a survey on the grounds that the target group is enormous has not done the arithmetic.
| Population (N) | Sample needed (n) |
|---|---|
| 100 people | 80 |
| 500 people | 218 |
| 1,000 people | 278 |
| 10,000 people | 370 |
| 100,000 people | 383 |
| 1,000,000 people | 385 |
| 340,000,000 people | 385 |
Calculated at 95 percent confidence, a 5 percent margin of error and p = 0.5.
More precision costs disproportionately many people
The margin of error is the most expensive input in the whole calculation, because it enters as a square. Moving from 5 to 3 percent costs almost three times as many participants, moving from 5 to 1 percent costs twenty-five times as many. For a thesis or capstone project, 95 percent confidence with a 5 percent margin is therefore the usual and often the only realistic setting.
| Margin of error | Confidence | Sample needed |
|---|---|---|
| 10 percent | 95 percent | 97 |
| 5 percent | 90 percent | 271 |
| 5 percent | 95 percent | 385 |
| 5 percent | 99 percent | 664 |
| 3 percent | 95 percent | 1,066 |
| 1 percent | 95 percent | 9,513 |
Calculated for a population of one million people and p = 0.5.
Qualitative research counts saturation, not numbers
For interviews and open formats the formula is useless, because nothing is being estimated as a proportion. The benchmark is theoretical saturation: you keep interviewing until new conversations stop producing new themes. In student projects that often means eight to fifteen interviews, though the justification in the text matters more than the count. The difference between the two routes is explained in the guide to qualitative and quantitative research methods.
Describing your sample in a thesis or paper
Your sample belongs in the methods section, immediately after the choice of research method and before the questionnaire itself. Nobody expects a perfect selection, only a traceable one: readers want to know who you studied, how you reached those people, and what limits follow for your results.
Points are quietly lost at exactly this stage. A convenience sample costs nothing, an undisclosed method costs a great deal, because reviewers then have to work out for themselves how solid your figures are. Naming the limits yourself reads as confidence rather than weakness. Check your program requirements as well, since research involving people usually needs approval from an Institutional Review Board before any data is collected.
What belongs in the methods section
- The population, named and clearly bounded.
- The sampling method you chose, with a reason.
- The route by which you reached participants.
- The fielding period, with start and end dates.
- The number of responses, split into started and completed.
- The composition by the characteristics that matter to your question.
- The limits that follow from the method.
One sentence covers the last line, for example: “The sample is a convenience sample, so the findings apply to the group surveyed and cannot be generalized to all students.” How the methods section fits together as a whole is set out in the guide to writing an empirical paper.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeCommon sampling mistakes
The most frequent mistakes happen before and after the arithmetic rather than inside it: when the population is defined, when the survey is distributed, and when the findings are written up. Four cases keep coming back.
All four are avoidable if you write down, before you start, exactly who your project is meant to say something about. Without that sentence on paper the target group drifts during fielding, usually toward whoever is easiest to reach.
The population stays vague
“Young people” is not a population, because no selection follows from it. Without an age range, a region and a time frame you can neither draw a sample nor check later whether it fits. The consequence only surfaces during analysis, when it stays unclear who the numbers actually describe. Write the population out as a full sentence before the first question exists.
Sample size is set on a hunch
Round numbers such as 100 or 200 look solid but are rarely justified. With a population of 500 people, 218 responses cover a 5 percent margin of error; with 100,000 people it takes 383. Setting the number in advance instead of calculating it means aiming too high and wasting time, or too low and losing the ability to say anything.
Response rate is left out of the plan
A large gap regularly sits between invitations sent and questionnaires completed. Needing 373 responses and contacting exactly 373 people ends well below target. Plan the distribution more generously and watch the response while the survey is still in the field, rather than hoping for a number at the end. A second, friendly reminder usually achieves more than extending the deadline by two weeks.
A skewed selection turns into a general claim
The most expensive mistake usually appears in the conclusion. A survey of 200 classmates becomes a sentence about “students in the United States”. Specialists spot the leap immediately, and it devalues the parts of the work that were done properly. Always phrase findings together with the group you actually surveyed.
Conclusion
Your sample decides the quality of your findings long before any analysis does. Choose the method deliberately, calculate the size once and properly, and write both into the methods section along with their limits. A small sample described honestly carries a project further than a large one nobody can reconstruct.
Where to go next
- Ready to build the questionnaire? Creating a questionnaire
- Unsure which level of measurement your questions use? Levels of measurement
- Already planning the analysis? Methods of analysis
Need 385 responses and only three weeks to get them?
With empirio.ai, an online survey tool from Germany, you build your survey in a few steps, share it by link and watch responses arrive in real time, so you can act while the survey is still in the field.
Frequently asked questions
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for free