Your questionnaire is back and 87 rows are sitting in a spreadsheet. Now what? Most people reach straight for the average without first checking whether these answers allow an average at all.
Descriptive statistics summarizes the cases you actually collected, using measures such as the mean, the median and the standard deviation. Inferential statistics goes one step further and uses that sample to estimate what is probably true of the whole population, together with a statement of how uncertain the estimate is. By the end you will know which measure your data allow and when you can skip the second step entirely.
📌 Key points at a glance
- Descriptive statistics describes the sample, inferential statistics estimates the population.
- The mode works at every level of measurement, the mean only from interval upward.
- Describing only your own group needs no significance test.
- An outlier drags the mean noticeably, the median barely moves.
- The Census Bureau publishes margins of error at 90 percent confidence.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
What is descriptive statistics?
Descriptive statistics is the branch of statistics that summarizes and presents an existing data set without going beyond it. Every statement applies to the cases you collected, and to not a single case beyond them.
That restriction sounds modest, yet it carries half the analysis. Nobody learns anything from 200 rows of raw data. Only when those rows become a handful of figures and one chart does a picture emerge of where the middle lies, how widely the answers spread and whether two variables move together. Those three questions are exactly what the three groups of measures answer.
Presentation belongs to descriptive statistics as well: frequency tables, bar and pie charts, histograms, box plots. The form follows the point you want to make, not the first option in your software menu. Nicola Döring places data analysis in the same order in her textbook Forschungsmethoden und Evaluation in den Sozial- und Humanwissenschaften (6th edition, Springer 2023, chapter 12): describe first, test second. For the wider process, see the overview of empirical research.
Descriptive and inferential statistics: the difference
The difference between descriptive and inferential statistics is a question of reach. Descriptive statistics only claims something about the people who answered. Inferential statistics claims something about everyone those people are meant to represent, and quantifies its own uncertainty while doing so.
Inferential statistics is therefore also called inductive statistics. It does two things: it estimates population values, usually as a confidence interval, and it tests hypotheses about differences or relationships. Any data set can be described. Only a sample drawn by a proper selection procedure supports inference.
Both halves depend on each other more than they first appear to. Without a clean description you will not notice that a distribution is heavily skewed, that a single value distorts everything, or that a comparison group contains three people. All three findings decide which test is admissible at all. That is why description opens the results section.
The measures used in descriptive statistics
The measures used in descriptive statistics fall into three groups: central tendency for the middle, dispersion for the spread, association for the relationship between two variables. Reporting only one group leaves half the analysis undone.
One example runs through this chapter. You asked 15 people in your program how many hours a week they spend writing their thesis. Sorted, the answers are 2, 3, 3, 4, 4, 4, 5, 5, 6, 6, 7, 8, 9, 10, 20. One person stands out, and that single answer demonstrates every measure.
Central tendency: mode, median and mean
Measures of central tendency say where the middle of a distribution lies. The mode is the most frequent value, here 4 hours, which appears three times. If two values occur equally often there are two modes, and if every value differs the mode says nothing at all. The median sits in the middle of the ordered series, so with 15 values it is the eighth: 5 hours. The mean is the total divided by the count, here 96 divided by 15, so 6.4 hours.
Three measures, three different answers to one question. The value of 20 is responsible. Remove it and the mean falls from 6.4 to 5.4 hours while the median stays at 5. The median is robust against outliers, the mean is not. With skewed variables such as income, wait times or workloads you should report both.
Dispersion: range, variance and standard deviation
Measures of dispersion say how far apart the values lie. The range is the distance between the largest and the smallest value, here 20 minus 2, so 18 hours. The variance gathers up the deviations from the mean: you square each deviation, add them all together and divide by n minus 1, which gives 19.4 here. Subtracting one is the usual route as long as your data are a sample and not the whole population. The standard deviation is the square root of the variance, roughly 4.4 hours.
Variance and standard deviation are often confused, although the distinction is simple: the variance is expressed in squared units and cannot be read intuitively, the standard deviation is expressed in hours and belongs next to the mean. The outlier pulls hard here too. Without the value of 20 the standard deviation drops from 4.4 to 2.4 hours.
Association: which coefficient suits which data
Measures of association describe how strongly two variables move together. The level of measurement decides which coefficient applies: Pearson's r for two metric variables, Spearman's rank correlation as soon as one of them is merely ordinal, Cramér's V for two nominal variables. Pearson and Spearman run between minus 1 and plus 1, Cramér's V between 0 and 1. The contingency coefficient, also widely used, reaches 1 only in its corrected form, which makes Cramér's V the safer choice.
⚠️ Careful
A high correlation coefficient proves no causal link. It only shows that two variables vary together. Write “is associated with” rather than “leads to” unless your design can establish a direction of effect. Pearson's r also detects straight-line relationships only, so a value near zero does not rule out a strong but curved relationship. Always plot a scatter diagram before you read a correlation.
Which measure may you use at which level of measurement?
What is permitted depends on the level of measurement of your variable, not on whether your software returns a number. Excel will happily calculate a mean for a variable in which 1 stands for Austin, 2 for Denver and 3 for Portland. The result is still meaningless.
The classification goes back to the psychologist Stanley Smith Stevens, who distinguished four types of scale in Science in 1946 and set out which measures remain meaningful for each (Stevens 1946). The table below gives the reading in common use today, which extends Stevens in places. Statisticians have criticized the scheme since the 1990s: read as a rigid rule, the four types mislead in borderline cases.
| Level of measurement | Permissible central tendency | Permissible dispersion |
|---|---|---|
| Nominal scale | Mode | none in the strict sense, report frequencies and proportions |
| Ordinal scale | Mode, median | Quartiles, percentiles |
| Interval scale | plus the arithmetic mean | Range, variance, standard deviation |
| Ratio scale | plus the geometric mean | plus the coefficient of variation |
Read the table cumulatively: whatever a lower level permits stays permissible at every higher one. The mode therefore works everywhere, the mean only from the interval scale upward. To work out the level of your own variables, see the article on the level of measurement.
The recurring argument concerns agreement scales running from “strongly disagree” to “strongly agree”. Strictly they are ordinal, because nobody can show that all respondents perceive the gaps between the options as equal. Research practice nevertheless often treats such items as interval data, especially where several items are combined into one scale. That is defensible only with symmetrical, fully labeled options, and the decision belongs in your methods section with a justification. The article on the Likert scale covers it in full.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
When you need no inferential statistics at all
Inferential statistics becomes unnecessary the moment nobody wants to generalize to a larger group. A significance test exists purely to move from a sample to a population. Remove that purpose and the reason for the test disappears with it.
The clearest case is a census of your own population. If all 42 members of your student organization answer and you want to say something about exactly those 42 people, the population is in front of you and there is nothing left to infer. That 19 of 42 prefer the later meeting slot is not an estimate with uncertainty attached, it is simply the result. If only some of them answer, it is no longer a census but a group that selected itself.
One limit applies nonetheless, and it belongs in your methods section. The rule holds as long as your claim is confined to exactly this group at exactly this moment. As soon as you read the result as evidence of something underlying it, future members or a cause, you are working with a model again and confidence intervals are back on the table. The methods literature files this case under superpopulation.
💡 Tip
Put one sentence in your methods section explaining why no test follows: “As this is a census of the population, no inferential procedures were applied.” At the defense that reads far more confidently than a test with no target.
The second case is the convenience sample. Sharing the questionnaire across three WhatsApp groups and one Instagram profile is not random selection, it is asking whoever was reachable. A test can still be calculated on such data, but its result does not travel beyond the respondents. That limitation belongs openly in your limitations section, and that is precisely where credit is earned. What makes a sample methodologically sound and when a survey counts as representative is covered in the two linked articles.
What the p-value says and what it does not
The p-value states how likely a result like yours, or a more extreme one, would be if the null hypothesis were true and every other assumption in your model held, random selection among them. It is a statement about the data under those assumptions, not about the assumptions themselves. A small p-value can therefore also mean that one of those assumptions is wrong.
The American Statistical Association published six principles on p-values in 2016 because misinterpretations had become widespread in the research literature. Principle two leaves no room: “P-values do not measure the probability that the studied hypothesis is true …” (Wasserstein and Lazar 2016). The sentence continues by ruling out the second common misreading: a p-value says just as little about how likely it is that your result came about by chance alone. A p-value of 0.03 does not mean your hypothesis is 97 percent likely. What your own two groups produce is shown by the p-value calculator, though the interpretation still remains your job.
Nor does a p-value measure the size or the importance of an effect, which is principle five of the same statement. With a very large sample even a tiny, practically irrelevant difference can become significant. Report an effect size alongside every p-value, and a confidence interval wherever possible.
How official statistics handles uncertainty
The U.S. Census Bureau shows what taking uncertainty seriously looks like in practice. It publishes a margin of error with every American Community Survey estimate, and it does so at the 90 percent confidence level rather than the 95 percent level common in academic work (Census Bureau, Understanding Error and Determining Statistical Significance). Two consequences follow. A figure without a statement about its precision is not a finished figure, and confidence intervals are only comparable at the same confidence level. For formulating your assumptions beforehand, see the article on developing hypotheses.
Descriptive statistics in your results section
The results section opens with descriptive statistics, and specifically with a description of the sample: number of cases, response rate, distribution by age, gender, year of study or whatever your research question requires. Only then come the measures for the substantive variables, and finally the hypothesis tests.
A simple rule of thumb governs presentation. Anything expressible in two sentences stays in the text. Anything covering more than three values goes into a table. Anything meant to show a distribution or a trend becomes a chart. Every table reports n and the measures that suit the level of measurement, so mean and standard deviation for metric variables and frequencies for nominal ones. Every chart has labeled axes with units.
What to calculate with
A spreadsheet is entirely sufficient for central tendency and dispersion. SPSS is standard at many universities but requires a license; JASP and jamovi are free, work in a similar way and cover the procedures a master's thesis usually calls for. R and Python are worth it if you already use them. Your grade depends on a transparent calculation, not on the software.
Things get considerably easier when the data arrive already structured. Online surveys let you download responses as an Excel or CSV file, which saves you retyping and the typing errors that come with it. empirio.ai, an online survey tool from Germany, shows frequencies and means directly in its results view and also exports the raw data as a file for your own analysis.
Common mistakes in the analysis
Most weaknesses in a results section are not arithmetic errors but matching errors: a measure is applied to data it was never meant for, or a figure is read as saying more than it can. Four patterns come up again and again.
All four go unnoticed while you write and stand out immediately when someone reads. Going through the results section for them before submission takes a quarter of an hour, and they are exactly what a committee asks about.
A mean for a nominal variable
A mean of 1.6 for a variable whose codes stand for majors is not information, it is a by-product of the coding. For nominal variables you report frequencies and proportions, nothing else. The test is quick: if the statement changes when you reassign the codes, the measure is inadmissible.
Reporting only the mean
A mean without a measure of dispersion withholds half the information. Some 6.4 hours a week can mean that everyone works about six hours, or that some respondents work two hours and others eleven. Standard deviation and sample size therefore belong on the same line as the mean.
Percentages without a base
“60 percent of respondents agreed” sounds solid until you learn that five people answered. Below roughly twenty cases a single person is worth more than five percentage points, which makes the absolute figure the more honest one. Always give it, so “three out of five”.
Confusing significance with relevance
A significant difference is not automatically an important difference. Whether an effect matters in practice is shown by the effect size, set against what is usual in your field and judged by what would actually make a difference in your case. Phrases such as “highly significant” imply a magnitude that the p-value simply does not measure.
Conclusion
Descriptive statistics is not a warm-up for the real analysis. In many student projects it is the whole analysis. Anyone who surveyed everybody, or who could not draw a random sample, should describe carefully and state the limits rather than run a test their data cannot carry. And anyone who does test should place an effect size next to the p-value.
An honestly described sample is worth more than an asterisk next to a number.
Where to go next
- Not sure which procedure fits? Methods of analysis at a glance
- Planning the empirical part? Writing an empirical thesis
- Unsure about your scales? Determining the level of measurement
- Want the whole process from the start? Quantitative research
Would you rather not retype your data?
With empirio.ai you build your survey free of charge and get frequencies, means and a raw data export without an extra step. Create a survey
