empirio.ai
Product
PricingLearn
  1. Learn
  2. /
  3. Empirical Research
  4. /
  5. Sample: How to Choose Your Method and Sample Size

Sample: How to Choose Your Method and Sample Size

How many answers do you really need? We go through the sampling methods, work the sample size out on an example and say plainly when the formula stops helping you.

Author at empirio.ai - Maria Malzewby Maria MalzewUpdated August 31, 2026Reading time 12 min

Summarize with

ChatGPTClaudePerplexityGoogle AI ModeGrok

In 1936 an American magazine mailed more than ten million ballots to predict a presidential election. Over two million people filled one out, and the forecast still named the wrong winner, missing by roughly 20 percentage points.

You determine a sample in three steps: define the population, choose a sampling method, and calculate the size that follows from both. Only random sampling lets you generalize to the wider population, and the sample size formula holds for random sampling alone. By the end you will know how many participants your project needs and how to justify that choice in your methods section.


📌 The key points at a glance

  • A sample is the subset of a larger population you actually study.
  • Only random sampling supports conclusions about the wider population.
  • At 95 percent confidence and a 5 percent margin, 385 people suffice.
  • Above one million people, population size no longer changes the answer.
  • The American Community Survey samples about 3.5 million addresses yearly.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.

Start for free

What is a sample?

A sample is a subset drawn from a larger group, the population, and studied on behalf of that group. Findings from the subset are generalized to the whole group, provided the subset was selected through a recognized statistical procedure rather than by convenience.

The population covers everyone or everything you want to make a statement about. Statistics writes the population as capital N and the sample as lowercase n. Studying every element of the population is a census. Studying part of it is a sample survey, and in practice that is almost always the only workable route.

Official statistics show how small a sample can be. The U.S. Census Bureau randomly selects about 3.5 million addresses each year for the American Community Survey, a fraction of all addresses in the country, and responding is required by law under Title 13 of the U.S. Code (as of August 2026). From that fraction come the figures on income, commuting and housing that guide federal funding decisions. Your own empirical research works on the same principle, on a smaller scale.

A census is the exception. A justified selection is the normal case.

Why a large sample is not automatically representative

A large sample is not automatically representative, because the number of responses says nothing about how those responses came about. If the selection method distorts the composition, every additional response only measures the wrong result more precisely.

The best known evidence comes from the magazine The Literary Digest. Before the 1936 election it mailed over ten million ballots, mostly to addresses taken from telephone directories and automobile registration lists. More than 2.3 million came back, a response rate just under a quarter. The forecast gave 55 percent to Alf Landon and 41 percent to Franklin D. Roosevelt. Roosevelt won, with 60.8 percent of the popular vote (The American Presidency Project, University of California Santa Barbara).

Almost every textbook offers the same explanation: only wealthier households owned a telephone and a car, so the selection was skewed. The case is less tidy than that. Peverill Squire analyzed a follow-up survey in Public Opinion Quarterly in 1988 and found that even people who owned both a car and a telephone voted mainly for Roosevelt. The second error sat in the response: Landon supporters returned the ballot far more often. How the total error splits between the two causes is still disputed. Squire (1988) weights the address lists more heavily, Dominic Lusinchi (2012) the response.

One uncomfortable rule follows for your own survey: who replies shapes the result as strongly as who was invited. Whether a survey ends up representative therefore depends on the method and the response, not on hitting a minimum number.

Sampling methods: the three routes to your selection

Sampling methods fall into three groups: arbitrary selection, purposive selection and random selection. Only random selection supports a statistically grounded conclusion about the population, because only there does every element have a known chance of being drawn.

Which group applies to you is rarely settled by the ideal and almost always by a practical question: do you have a complete list of your population at all? No such list exists for “everyone who bikes in Chicago”, so genuine random selection is off the table. The honest answer in your methods section is then that you selected purposively or arbitrarily.

MethodHow people are selectedWhen it fits
Arbitrary selectionwhoever happens to be reachableearly exploration, pilot testing
Purposive selectionby criteria set in advanceextreme cases, rare groups
Random selectionby a random procedure from a listconclusions about the population

The three most common random samples

Within random selection the methods literature distinguishes three variants that differ sharply in effort. All three meet the same condition, that chance decides who takes part. They simply apply that chance at different points in the procedure, and that changes how precise the final result becomes.

  • Simple random sample: every element of the population is drawn directly and with equal probability.
  • Stratified random sample: the population is first split by a characteristic, then drawn within each stratum.
  • Cluster sample: whole groups such as school classes are drawn, and everyone inside is surveyed.

For student projects the stratified variant is often the most sensible, because it keeps an important group from ending up too small by pure chance. How to reach enough people for the draw is covered in the guide to finding survey participants.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.

Start for free

How to calculate sample size: formula, values and an example

Sample size follows from four inputs: the size of the population, the margin of error you accept, the confidence level you want, and the expected spread of answers. Together they give the minimum number of people who need to complete your survey.

The four inputs in detail:

  • Population (N): everyone you want to make a statement about.
  • Margin of error (e): how far your result may sit from the true value, commonly 5 percent.
  • Confidence level: how reliably that interval captures the true value, commonly 95 percent.
  • z-score: the confidence level expressed as a number, 1.645 at 90 percent, 1.96 at 95 percent, 2.576 at 99 percent.
  • Spread (p): the expected distribution of answers, set to 0.5 without prior knowledge, the least favorable case.

Two steps turn those inputs into a number. The first gives the value for a very large population, the second corrects it downward when your group is small:

n0 = z² × p × (1 − p) / e²
n  = n0 / (1 + n0 / N)

A worked example: you survey the students at your university, 12,000 people in total. At 95 percent confidence (z = 1.96), a 5 percent margin of error (e = 0.05) and p = 0.5, the first step gives 384.16 and the second 372.2. Rounded up, you need 373 completed responses. Because a share of respondents always drops out partway, send the questionnaire to considerably more people than that.

When the sample size formula does not apply

The formula assumes random selection. If you share your survey through an Instagram story, a group chat or a bulletin board in your department, chance is not deciding who takes part. Willingness among the people who already see you is. No margin of error can be calculated for a convenience sample of that kind, not even with 500 responses.

⚠️ Careful

385 responses from an Instagram story say nothing about the United States, only something about your reach. Do not put a margin of error in your paper in cases like that. Name the method instead, limit the claim to the group you actually reached, and explain why that is enough for your research question.

How many participants do you really need?

For most projects the sample you need sits between 80 and 400 people. At 95 percent confidence and a 5 percent margin of error it is 385 people once the population passes one million, and considerably fewer as soon as your group is manageable.

The surprising part is how little population size matters. Between one million and 340 million people the required sample does not change at all. Anyone who inflates a survey on the grounds that the target group is enormous has not done the arithmetic.

Population (N)Sample needed (n)
100 people80
500 people218
1,000 people278
10,000 people370
100,000 people383
1,000,000 people385
340,000,000 people385

Calculated at 95 percent confidence, a 5 percent margin of error and p = 0.5.

More precision costs disproportionately many people

The margin of error is the most expensive input in the whole calculation, because it enters as a square. Moving from 5 to 3 percent costs almost three times as many participants, moving from 5 to 1 percent costs twenty-five times as many. For a thesis or capstone project, 95 percent confidence with a 5 percent margin is therefore the usual and often the only realistic setting.

Margin of errorConfidenceSample needed
10 percent95 percent97
5 percent90 percent271
5 percent95 percent385
5 percent99 percent664
3 percent95 percent1,066
1 percent95 percent9,513

Calculated for a population of one million people and p = 0.5.

Qualitative research counts saturation, not numbers

For interviews and open formats the formula is useless, because nothing is being estimated as a proportion. The benchmark is theoretical saturation: you keep interviewing until new conversations stop producing new themes. In student projects that often means eight to fifteen interviews, though the justification in the text matters more than the count. The difference between the two routes is explained in the guide to qualitative and quantitative research methods.

Describing your sample in a thesis or paper

Your sample belongs in the methods section, immediately after the choice of research method and before the questionnaire itself. Nobody expects a perfect selection, only a traceable one: readers want to know who you studied, how you reached those people, and what limits follow for your results.

Points are quietly lost at exactly this stage. A convenience sample costs nothing, an undisclosed method costs a great deal, because reviewers then have to work out for themselves how solid your figures are. Naming the limits yourself reads as confidence rather than weakness. Check your program requirements as well, since research involving people usually needs approval from an Institutional Review Board before any data is collected.

What belongs in the methods section

  • The population, named and clearly bounded.
  • The sampling method you chose, with a reason.
  • The route by which you reached participants.
  • The fielding period, with start and end dates.
  • The number of responses, split into started and completed.
  • The composition by the characteristics that matter to your question.
  • The limits that follow from the method.

One sentence covers the last line, for example: “The sample is a convenience sample, so the findings apply to the group surveyed and cannot be generalized to all students.” How the methods section fits together as a whole is set out in the guide to writing an empirical paper.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.

Start for free

Common sampling mistakes

The most frequent mistakes happen before and after the arithmetic rather than inside it: when the population is defined, when the survey is distributed, and when the findings are written up. Four cases keep coming back.

All four are avoidable if you write down, before you start, exactly who your project is meant to say something about. Without that sentence on paper the target group drifts during fielding, usually toward whoever is easiest to reach.

The population stays vague

“Young people” is not a population, because no selection follows from it. Without an age range, a region and a time frame you can neither draw a sample nor check later whether it fits. The consequence only surfaces during analysis, when it stays unclear who the numbers actually describe. Write the population out as a full sentence before the first question exists.

Sample size is set on a hunch

Round numbers such as 100 or 200 look solid but are rarely justified. With a population of 500 people, 218 responses cover a 5 percent margin of error; with 100,000 people it takes 383. Setting the number in advance instead of calculating it means aiming too high and wasting time, or too low and losing the ability to say anything.

Response rate is left out of the plan

A large gap regularly sits between invitations sent and questionnaires completed. Needing 373 responses and contacting exactly 373 people ends well below target. Plan the distribution more generously and watch the response while the survey is still in the field, rather than hoping for a number at the end. A second, friendly reminder usually achieves more than extending the deadline by two weeks.

A skewed selection turns into a general claim

The most expensive mistake usually appears in the conclusion. A survey of 200 classmates becomes a sentence about “students in the United States”. Specialists spot the leap immediately, and it devalues the parts of the work that were done properly. Always phrase findings together with the group you actually surveyed.

Conclusion

Your sample decides the quality of your findings long before any analysis does. Choose the method deliberately, calculate the size once and properly, and write both into the methods section along with their limits. A small sample described honestly carries a project further than a large one nobody can reconstruct.

Where to go next

  • Ready to build the questionnaire? Creating a questionnaire
  • Unsure which level of measurement your questions use? Levels of measurement
  • Already planning the analysis? Methods of analysis

Need 385 responses and only three weeks to get them?

With empirio.ai, an online survey tool from Germany, you build your survey in a few steps, share it by link and watch responses arrive in real time, so you can act while the survey is still in the field.

Create a survey for free

Frequently asked questions

At 95 percent confidence and a 5 percent margin of error, a sample needs 385 completed responses once the population passes one million people. Smaller groups need far fewer: 383 at 100,000 people, 370 at 10,000, 278 at 1,000 and 80 at 100. All of these figures assume random selection.

A sample is representative when it was drawn by a random procedure from a fully listed population and the response is not systematically skewed. Size alone does not decide it. The Literary Digest survey of 1936 collected more than 2.3 million responses and still predicted the wrong winner of the presidential election.

A population covers everyone or everything a statement is meant to apply to, while a sample covers only the subset actually studied. Statistics writes the population as capital N and the sample as lowercase n. When every element is studied, the exercise is a census rather than a sample survey.

Sampling methods fall into arbitrary selection, purposive selection and random selection. Random selection has three common variants: the simple random sample, the stratified random sample and the cluster sample. Only random selection supports a statistically grounded conclusion about the population, because every element has a known chance of being drawn.

Sample size follows from four inputs: population N, margin of error e, the z-score for the confidence level, and the expected spread p. First calculate n0 = z² × p × (1 − p) / e², then n = n0 / (1 + n0 / N). With N = 12,000, z = 1.96, e = 0.05 and p = 0.5 the result is 373 people.

The U.S. Census Bureau randomly selects about 3.5 million addresses each year for the American Community Survey. Responding is required by law under Title 13 of the U.S. Code, which is why the survey reaches response rates that voluntary surveys rarely approach (as of August 2026).

Capital N stands for the population, meaning everyone or everything a statement is meant to apply to. Lowercase n stands for the sample, meaning the number of cases actually studied. In results reporting, n often also marks the number of cases in a single subgroup, such as n = 42 for one age band.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.

Start for free

Related articles

Empirical Research

Fundamentals of Empirical Research | empirio

What is empirical research and what is it used for in scientific research?

Read more
Survey Tips

When is my online survey representative?

A survey does not turn representative because a lot of people answered. It turns representative through who you select. We explain the criteria and what is reachable online.

Read more
Empirical Research

Quantitative Research: Process, Methods & Example

In this article, we examine quantitative research and which subject areas it is best suited for. Using a concrete example - a quantitative survey - we also show you how to implement this research method in your thesis.

Read more

Create a survey

With empirio.ai you can create a modern online survey in minutes.

Start for free

Footer

empirio.ai

Create surveys in no time - with AI, modern templates and data protection from Germany.

Sign up free

Features

  • Product
  • Pricing
  • Integrations (Docs)
  • MCP Server
  • Create surveys with ChatGPT
  • Create surveys with Claude
  • AI info

Solutions

  • For Schools
  • For Universities
  • For Companies
  • For Students
  • Create Employee Surveys
  • Create Customer Surveys
  • All Solutions

Templates

  • Employee Surveys
  • Customer Feedback
  • Market Research
  • HR & Recruiting
  • Events
  • Education
  • All Templates

Tools & Help

  • Help
  • Learn
  • Glossary
  • AI tools for academic writing
  • Plagiarism checker

Legal

  • Privacy
  • Terms
  • Imprint
  • About us
  • United States
  • United Kingdom
  • Germany
  • Austria
  • Switzerland
  • France
  • Spain
  • Italy
  • Netherlands

© 2026 empirio.ai

llms.txt