Two theses land on a committee table: thirty charts in one, a single mean and three interview quotes in the other. Which one was analyzed better? Not the one with more output, but the one whose analysis matches the question it set out to answer.
That is what data analysis methods are for, and it is where most projects go sideways. Data analysis methods are the systematic procedures you use to organize, condense, and test collected material so that it answers your research question. Three things decide which one applies: your research question, the kind of material you have, and for numbers the level of measurement. By the end of this article you will be able to defend your choice instead of asserting it.
📌 Key points at a glance
- The research question decides the method, not the sample size.
- Descriptive statistics always come before inferential ones.
- Level of measurement limits which statistic you may report.
- Qualitative material passes through recording, transcription, and coding.
- A p-value measures neither hypothesis probability nor effect size.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeWhat are data analysis methods?
Data analysis methods are defined, repeatable procedures for ordering, condensing, and interpreting collected material. They begin where data collection ends and finish where numbers or text have become a reasoned answer to the research question.
Several labels describe the same activity. Data analysis, analytic strategy, and analysis procedures are used interchangeably in American academic writing, and none is more precise than the others. A different distinction carries far more weight: the one between preparation, analysis, and interpretation.
Preparation makes the material usable at all, which means checking missing values, coding responses, and transcribing recordings. Analysis turns the prepared material into statistics or categories. Interpretation reads those results against the research question and the existing literature. Folding interpretation into analysis is the fastest way to lose the transparency a committee is looking for.
Choosing a data analysis method: three questions
Choosing a data analysis method means answering three questions in order: what do you want to find out, what kind of material do you have, and at what level of measurement was it recorded? Only the third question narrows the field far enough that a single technique remains.
The order cannot be reversed. Picking a technique first because it appeared in a course, then reshaping the research question around it, produces a study that computes correctly and explains nothing. In the United States the sequence also has a practical consequence: your IRB submission describes the planned analysis, so the decision is on record before the first response arrives.
| Question | What you check | What follows from it |
|---|---|---|
| Purpose | describing, comparing, or understanding | describing leads to statistics, understanding to text analysis |
| Type of material | numbers, or text, image, and audio | numbers allow statistics, text requires categories |
| Level of measurement | nominal, ordinal, interval, or ratio | limits which statistics and tests are permitted |
Write the three answers down before you open any software. Those lines become the opening paragraph of your methods section, and they head off the most common committee comment of all: that the analysis is competent but unmotivated.
Quantitative data analysis: describing and inferring
Quantitative data analysis falls into two groups: descriptive statistics summarize the sample in front of you, while inferential statistics ask whether a finding holds in the wider population. An empirical study needs both, because either one alone leaves the argument unfinished.
The order never changes. Descriptive analysis shows how your sample is composed and whether the values are plausible at all. Only after that is it worth asking whether a difference would survive outside your sample. The full comparison sits in our article on descriptive and inferential statistics.
Descriptive statistics: summarizing your own sample
Descriptive statistics condense your data into a handful of figures without generalizing beyond the sample. They cover frequencies and proportions, measures of central tendency such as the mean, median, and mode, measures of spread such as range, standard deviation, and variance, and measures of association such as the correlation coefficient, which runs from -1 to +1 with 0 indicating no linear relationship.
One error shows up again and again here: correlation does not establish cause. When two variables move together, an unmeasured third variable may be driving both. Causal claims require a design that rules out competing explanations, not a larger correlation coefficient.
Inferential statistics: from sample to population
Inferential statistics use significance tests to judge whether a difference found in a sample is plausible in the population. The standard tool is null hypothesis significance testing: you set an alternative hypothesis against a null hypothesis and calculate how well the data fit the null.
What that calculation does not do is settle the question for you. Confidence intervals, effect sizes, and the design of the study carry at least as much of the argument, and a test result that arrives without them is difficult to interpret in either direction.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeWhat a p-value does and does not tell you
A p-value indicates how incompatible your data are with a specified statistical model, and that is the whole of it. In 2016 the American Statistical Association published six principles on the use and interpretation of p-values, the first time the profession issued a formal statement on the question (American Statistical Association, 2016).
Three of those principles matter directly for a thesis. A p-value does not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. A p-value, or statistical significance, does not measure the size of an effect or the importance of a result. And scientific conclusions should not be based only on whether a p-value passes a specific threshold.
What to report instead of a bare p-value
The ASA points to approaches that emphasize estimation over testing, naming confidence, credibility, and prediction intervals, Bayesian methods, and alternative measures of evidence such as likelihood ratios or Bayes factors. For most student projects the practical version is simple: report the effect size and a confidence interval next to every test, and say in plain language what the difference means.
The statement also names the behaviors that follow from treating a threshold as a verdict. Ron Wasserstein, the ASA executive director, put it bluntly: the p-value was never intended to be a substitute for scientific reasoning. Then ASA president Jessica Utts pointed to the file-drawer effect, where studies without significant results never get published, and to the practices known as p-hacking and data dredging.
⚠️ Watch out
A test can reject a null hypothesis. It can never prove one. A result of p greater than 0.05 does not show that no difference exists, only that your data are not incompatible enough with the null to reject it. Writing “no effect was found” when you mean “no significant effect was detected” is the single most common misstatement in student results chapters.
Which statistics fit which level of measurement?
Level of measurement decides which statistic and which test your data will support. Nominal data allow frequencies only, ordinal data add rank statements and the median, and interval and ratio data are the first to permit the arithmetic mean and everything built on it.
The table below stays deliberately short and covers the cases that appear in almost every capstone, thesis, and dissertation project. Our article on levels of measurement explains what sits behind each step.
| Level of measurement | Statistics permitted | Typical techniques |
|---|---|---|
| Nominal | frequency, proportion, mode | cross-tabulation, chi-square test |
| Ordinal | median, quartiles, rank correlation | rank-based tests, Spearman correlation |
| Interval and ratio | mean, standard deviation, variance | t-test, analysis of variance, Pearson correlation |
This is not a formality. An arithmetic mean across nominal categories, such as an average major, produces a number with no meaning, and that is exactly the kind of number a committee member catches in seconds. Likert items are the recurring gray area: strictly ordinal, frequently treated as interval, and always worth a sentence of justification plus a median reported alongside the mean.
Qualitative data analysis: from interview to code
Qualitative data analysis condenses text, audio, or image material into an answer to the research question by way of a coding frame. Unlike numbers, this material does not arrive in analyzable form at all, which is why a preparation step always comes first.
The sequence is strikingly similar across approaches, even though the schools justify their procedures very differently. Three steps run through practically every qualitative analysis.
The three steps of any qualitative analysis
- Recording. Interviews, focus groups, or observations are recorded or logged, with documented consent and, in the United States, IRB approval obtained before collection begins.
- Transcription. The recording is written up according to rules set in advance. Whether accent, pauses, and filler words are captured depends on the approach and belongs in the methods section rather than being decided on the fly.
- Coding. The transcript is assigned to categories passage by passage. Codes are either developed inductively from the material or applied deductively from theory and the interview guide.
Approaches you are most likely to use
Qualitative content analysis and thematic analysis are the two coding-based approaches most American students meet first. Both move from familiarization through initial coding to categories or themes reviewed against the full data set. Philipp Mayring systematized the content analysis variant from the first edition of 1982 onward and distinguishes inductive category development from deductive category application.
Grounded theory, developed by Glaser and Strauss, takes a different route and aims at building theory from the material, alternating continuously between collecting and analyzing. Interpretive and sequence-based approaches work minutely through single passages and suit very small samples. Whichever you choose, report how codes were developed and, if a second coder was involved, report intercoder reliability.
Reading tip: Mayring, Philipp (2010). Qualitative Inhaltsanalyse. Grundlagen und Techniken. 12th edition. Weinheim and Basel: Beltz. An English summary by the same author from 2014 is freely available through the Social Science Open Access Repository.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for freeWriting the data analysis part of your methods section
The data analysis write-up belongs in the methods section and answers four questions: which technique you chose, why it fits the research question, how you applied it in practice, and which software you used. Four to six paragraphs are usually enough.
Reasoning beats description. A sentence such as “data were analyzed in SPSS” reveals nothing about a methodological decision. “Because the dependent variable is ordinal, a rank-based test was used in place of a t-test” shows that you know the limits of your material, and that is what a committee rewards.
Write the section so that another researcher could repeat your analysis from it. That is exactly what the criteria of objectivity, reliability, and validity are testing, and it is where projects with plenty of computation and little documentation fall apart. Your program handbook and your advisor set the binding requirements.
Common mistakes in data analysis
Most points lost in empirical projects go missing through four recurring errors of reasoning rather than through arithmetic. All four are avoidable if you know about them before you start computing.
Mistake 1: fixing the method before the research question
Choosing an analysis technique before the research question is settled builds the project backwards. The result is a set of correctly computed figures that add up to nothing. Formulate your hypotheses before you type a single line into your analysis software.
Mistake 2: computing statistics the measurement level will not support
A mean across grades, ranks, or answer categories looks harmless and remains indefensible. Check each variable separately for its level of measurement and record the result before you analyze anything. Twenty minutes of work spares you an awkward question at the defense.
Mistake 3: confusing significance with importance
A statistically significant result says nothing about whether the effect is large or practically relevant. In very large samples even a trivial difference reaches significance. Report an effect size alongside every p-value and explain what the difference means in substantive terms.
Mistake 4: planning the analysis after data collection
Fielding the questionnaire first and thinking about analysis afterwards usually reveals that a decisive question was asked in the wrong format. An open-ended question cannot be converted into a scale after the fact. Plan the analysis while you are still designing the questionnaire.
Conclusion
The right data analysis method is not found in a list. It follows from your research question, your type of material, and your level of measurement. Settling those three points before data collection spares you the uncomfortable discovery that the best data in the world answer the wrong question.
The data do not decide the method. The question decides both.
Where to go next
- Want the difference between the two branches of statistics? Descriptive and inferential statistics
- Unsure about the measurement level of your variables? Levels of measurement
- Looking for the whole research process in one place? Empirical research
Thinking about analysis while you build the survey?
With empirio.ai, an online survey tool from Germany, you see frequencies, cross-tabulations and filters in a dashboard that updates with every new response.
Frequently asked questions
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with 100% data protection from Germany.
Start for free