In 1936 an American magazine posted more than ten million ballot papers to predict a presidential election. Over two million people filled one in, and the forecast still named the wrong winner, by roughly 20 percentage points.
You work out a sample in three steps: define the population, choose a sampling method, and calculate the size that follows from both. Only random sampling lets you draw conclusions about the wider population, and the sample size formula only holds for random sampling. By the end you will know how many participants your project needs and how to justify the choice in your methodology.
📌 The key points at a glance
- A sample is the subset of a larger population you actually study.
- Only random sampling supports conclusions about the wider population.
- At 95 per cent confidence and a 5 per cent margin, 385 people suffice.
- Above one million people, population size no longer changes the answer.
- The ONS Labour Force Survey reached 62,630 people in early 2026.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
What is a sample?
A sample is a subset drawn from a larger group, the population, and studied on behalf of that group. Findings from the subset are generalised to the whole group, provided the subset was selected through a recognised statistical procedure rather than by convenience.
The population covers everyone or everything you want to make a statement about. Statistics writes the population as capital N and the sample as lower case n. Studying every element of the population is a census. Studying part of it is a sample survey, and in practice that is almost always the only workable route.
Official statistics show how small a sample can be. The Office for National Statistics reports that the Labour Force Survey reached 62,630 people in 28,655 households between April and June 2026, with a response rate of 23.4 per cent (as of August 2026). Taking part is voluntary, and the low response is one reason the ONS currently publishes the results as official statistics in development. Your own empirical research works on the same principle, on a smaller scale.
A census is the exception. A justified selection is the normal case.
Why a large sample is not automatically representative
A large sample is not automatically representative, because the number of responses says nothing about how those responses came about. If the selection method distorts the composition, every extra response only measures the wrong result more precisely.
The best known evidence comes from the magazine The Literary Digest. Before the 1936 United States election it posted over ten million ballot papers, mostly to addresses taken from telephone directories and vehicle registration lists. More than 2.3 million came back, a response rate just under a quarter. The forecast gave 55 per cent to Alf Landon and 41 per cent to Franklin D. Roosevelt. Roosevelt won, with 60.8 per cent of the popular vote (The American Presidency Project, University of California Santa Barbara).
Almost every textbook offers the same explanation: only wealthier households owned a telephone and a car, so the selection was skewed. The case is less tidy than that. Peverill Squire analysed a follow-up survey in Public Opinion Quarterly in 1988 and found that even people who owned both a car and a telephone voted mainly for Roosevelt. The second error sat in the response: Landon supporters returned the ballot far more often. How the total error splits between the two causes is still disputed. Squire (1988) weights the address lists more heavily, Dominic Lusinchi (2012) the response.
One uncomfortable rule follows for your own survey: who replies shapes the result as strongly as who was invited. Whether a survey ends up representative therefore depends on the method and the response, not on hitting a minimum number.
Sampling methods: the three routes to your selection
Sampling methods fall into three groups: arbitrary selection, purposive selection and random selection. Only random selection supports a statistically grounded conclusion about the population, because only there does every element have a known chance of being drawn.
Which group applies to you is rarely settled by the ideal and almost always by a practical question: do you have a complete list of your population at all? No such list exists for “everyone who cycles in Manchester”, so genuine random selection is off the table. The honest answer in your methodology is then that you selected purposively or arbitrarily.
| Method | How people are selected | When it fits |
|---|---|---|
| Arbitrary selection | whoever happens to be reachable | early exploration, pilot testing |
| Purposive selection | by criteria set in advance | extreme cases, rare groups |
| Random selection | by a random procedure from a list | conclusions about the population |
The three most common random samples
Within random selection the methods literature distinguishes three variants that differ sharply in effort. All three meet the same condition, that chance decides who takes part. They simply apply that chance at different points in the procedure, and that changes how precise the final result becomes.
- Simple random sample: every element of the population is drawn directly and with equal probability.
- Stratified random sample: the population is first split by a characteristic, then drawn within each stratum.
- Cluster sample: whole groups such as school classes are drawn, and everyone inside is surveyed.
For student projects the stratified variant is often the most sensible, because it stops an important group from ending up too small by pure chance. How to reach enough people for the draw is covered in the guide to finding survey participants.
How to calculate sample size: formula, values and an example
Sample size follows from four inputs: the size of the population, the margin of error you accept, the confidence level you want, and the expected spread of answers. Together they give the minimum number of people who need to complete your survey.
The four inputs in detail:
- Population (N): everyone you want to make a statement about.
- Margin of error (e): how far your result may sit from the true value, commonly 5 per cent.
- Confidence level: how reliably that interval captures the true value, commonly 95 per cent.
- z-score: the confidence level expressed as a number, 1.645 at 90 per cent, 1.96 at 95 per cent, 2.576 at 99 per cent.
- Spread (p): the expected distribution of answers, set to 0.5 without prior knowledge, the least favourable case.
Two steps turn those inputs into a number. The first gives the value for a very large population, the second corrects it downwards when your group is small:
n0 = z² × p × (1 − p) / e²
n = n0 / (1 + n0 / N)
A worked example: you survey the students at your university, 12,000 people in total. At 95 per cent confidence (z = 1.96), a 5 per cent margin of error (e = 0.05) and p = 0.5, the first step gives 384.16 and the second 372.2. Rounded up, you need 373 completed responses. The same calculation with your own figures is handled by the sample size calculator. Because a share of respondents always drops out part way, send the questionnaire to considerably more people than that.
When the sample size formula does not apply
The formula assumes random selection. If you share your survey through an Instagram story, a WhatsApp group or a noticeboard in your department, chance is not deciding who takes part. Willingness among the people who already see you is. No margin of error can be calculated for a convenience sample of that kind, not even with 500 responses.
⚠️ Careful
385 responses from an Instagram story say nothing about the United Kingdom, only something about your reach. Do not put a margin of error in your dissertation in cases like that. Name the method instead, limit the claim to the group you actually reached, and explain why that is enough for your research question.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
How many participants do you really need?
For most projects the sample you need sits between 80 and 400 people. At 95 per cent confidence and a 5 per cent margin of error it is 385 people once the population passes one million, and considerably fewer as soon as your group is manageable.
The surprising part is how little population size matters. Between one million and 69 million people the required sample does not change at all. Anyone who inflates a survey on the grounds that the target group is enormous has not done the arithmetic.
| Population (N) | Sample needed (n) |
|---|---|
| 100 people | 80 |
| 500 people | 218 |
| 1,000 people | 278 |
| 10,000 people | 370 |
| 100,000 people | 383 |
| 1,000,000 people | 385 |
| 69,000,000 people | 385 |
Calculated at 95 per cent confidence, a 5 per cent margin of error and p = 0.5.
More precision costs disproportionately many people
The margin of error is the most expensive input in the whole calculation, because it enters as a square. Moving from 5 to 3 per cent costs almost three times as many participants, moving from 5 to 1 per cent costs twenty-five times as many. The other way round, the margin of error calculator shows you how precise your result already is at a number of participants you have reached. For a dissertation, 95 per cent confidence with a 5 per cent margin is therefore the usual and often the only realistic setting.
| Margin of error | Confidence | Sample needed |
|---|---|---|
| 10 per cent | 95 per cent | 97 |
| 5 per cent | 90 per cent | 271 |
| 5 per cent | 95 per cent | 385 |
| 5 per cent | 99 per cent | 664 |
| 3 per cent | 95 per cent | 1,066 |
| 1 per cent | 95 per cent | 9,513 |
Calculated for a population of one million people and p = 0.5.
Qualitative research counts saturation, not numbers
For interviews and open formats the formula is useless, because nothing is being estimated as a proportion. The benchmark is theoretical saturation: you keep interviewing until new conversations stop producing new themes. In student projects that often means eight to fifteen interviews, though the justification in the text matters more than the count. The difference between the two routes is explained in the guide to qualitative and quantitative research methods.
Describing your sample in a dissertation
Your sample belongs in the methodology chapter, immediately after the choice of research method and before the questionnaire itself. Nobody expects a perfect selection, only a traceable one: markers want to read who you studied, how you reached those people, and what limits follow for your results.
Marks are quietly lost at exactly this point. A convenience sample costs nothing, an undisclosed method costs a great deal, because markers then have to work out for themselves how solid your figures are. Naming the limits yourself reads as confidence rather than weakness. Check your module handbook as well, since most UK universities also require ethics approval before you collect any data.
What belongs in the methodology chapter
- The population, named and clearly bounded.
- The sampling method you chose, with a reason.
- The route by which you reached participants.
- The fieldwork period, with start and end dates.
- The number of responses, split into started and completed.
- The composition by the characteristics that matter to your question.
- The limits that follow from the method.
One sentence covers the last line, for example: “The sample is a convenience sample, so the findings apply to the group surveyed and cannot be generalised to all students.” How the methodology chapter fits together as a whole is set out in the guide to writing an empirical paper.
Common sampling mistakes
The most frequent mistakes happen before and after the arithmetic rather than inside it: when the population is defined, when the survey is distributed, and when the findings are written up. Four cases keep coming back.
All four are avoidable if you write down, before you start, exactly who your project is meant to say something about. Without that sentence on paper the target group drifts during fieldwork, usually towards whoever is easiest to reach.
The population stays vague
“Young people” is not a population, because no selection follows from it. Without an age range, a region and a time frame you can neither draw a sample nor check later whether it fits. The consequence only surfaces during analysis, when it stays unclear who the numbers actually describe. Write the population out as a full sentence before the first question exists.
Sample size is set on a hunch
Round numbers such as 100 or 200 look solid but are rarely justified. With a population of 500 people, 218 responses cover a 5 per cent margin of error; with 100,000 people it takes 383. Setting the number in advance instead of calculating it means aiming too high and wasting time, or too low and losing the ability to say anything.
Response rate is left out of the plan
A large gap regularly sits between invitations sent and questionnaires completed. Needing 373 responses and contacting exactly 373 people ends well below target. Plan the distribution more generously and watch the response during fieldwork rather than hoping for a number at the end. How many invitations you need can be worked out beforehand if you calculate the response rate. A second, friendly reminder usually achieves more than extending the deadline by a fortnight.
A skewed selection turns into a general claim
The most expensive mistake usually appears in the conclusion. A survey of 200 coursemates becomes a sentence about “students in the UK”. Specialists spot the leap immediately, and it devalues the parts of the work that were done properly. Always phrase findings together with the group you actually surveyed.
Conclusion
Your sample decides the quality of your findings long before any analysis does. Choose the method deliberately, calculate the size once and properly, and write both into the methodology chapter along with their limits. A small sample described honestly carries a project further than a large one nobody can reconstruct.
Where to go next
- Ready to build the questionnaire? Creating a questionnaire
- Unsure which level of measurement your questions use? Levels of measurement
- Already planning the analysis? Methods of analysis
Need 385 responses and only three weeks to get them?
With empirio.ai, an online survey tool from Germany, you build your survey in a few steps, share it by link and watch responses arrive in real time, so you can act while the field is still open.
