Two students code the same twelve interviews and arrive at two different results. Which of the two analyses is the better one gets decided by quality criteria, and qualitative research applies different ones than quantitative research.
Quality criteria are the standards that show whether a study is scientifically usable. Quantitative research in the United States rests on validity and reliability, joined by fairness in the current testing standards. Qualitative research has no single catalog, only competing proposals, the best known from Lincoln and Guba. After this guide you know which catalog fits your method, how to justify the choice and what to do when a German source names a third criterion.
📌 The key points
- Quantitative work in the US rests on validity and reliability.
- Fairness joins them as a third foundation in the 2014 standards.
- Lincoln and Guba name four criteria under the label trustworthiness.
- Sarah J. Tracy proposes eight big-tent criteria for qualitative work.
- German sources add a third criterion called Objektivität.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
What quality criteria are and why the answer depends on the tradition
Quality criteria are the benchmarks by which the quality of a study can be judged. They answer a single question: does a result come from the procedure that was used, or from the person who used it?
The difference between an everyday observation and a scientific statement sits exactly here. “I noticed that the mood on the team has soured” is a perception. Only when it becomes traceable how that statement came about, with which instrument, from which people and under which rules, does it turn into a finding that colleagues can argue about. Quality criteria are the language for that difference.
Which criteria apply, however, depends on the tradition you work in. Quantitative designs judge an instrument, so they ask about measurement error and about the construct behind a score. Qualitative designs judge a research process, so they ask about documentation, reasoning and traceability. Reading mostly American literature also pushes you toward different names than reading mostly German literature.
A quality criterion is a scholarly convention, not a rule of law. No statute and no examination code prescribes which criteria a thesis has to check, and graders rarely reward meeting all of them. Graders reward being able to name which ones you met and where your data ran into limits. Your department, your advisor and, for studies involving people, your institutional review board set the binding expectations. Where the section belongs in your outline is equally open, and both common placements are weighed up in Objectivity, Reliability, Validity: Definition and Example.
Quality criteria in quantitative research: validity and reliability
Quantitative research in the United States is judged on validity and reliability, and the current testing standards place fairness beside them as a third foundation.
Authoritative for that framing are the Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education. Part I of the 2014 edition is titled Foundations and holds exactly three chapters: Validity, Reliability/Precision and Errors of Measurement, and Fairness in Testing. A chapter on objectivity does not exist there.
| Foundation | The question behind it | What you point to |
|---|---|---|
| Validity | Does the score mean what you claim? | evidence from content, structure and relations |
| Reliability | How large is the measurement error? | coefficients across items, occasions and raters |
| Fairness | Does the instrument work for every group? | evidence across subgroups, accessible design |
Ranking the three is tempting and misleading. An unreliable measurement cannot be valid, because random noise carries no meaning. Reliability on its own guarantees nothing either: a scale that reads two pounds too high measures reliably and still measures wrong. How each of the three gets examined in practice is set out in Objectivity, Reliability, Validity: Definition and Example.
Types of validity evidence, and why the classic labels were retired
Most textbooks still teach content validity, construct validity and criterion validity as three kinds of validity. The 2014 Standards dropped that vocabulary on purpose. Validity is treated there as a unitary concept, and the text speaks of types of validity evidence rather than distinct types of validity, avoiding terms such as content validity or predictive validity.
Five sources of evidence replace the old labels: evidence based on test content, on response processes, on internal structure, on relations to other variables and on the consequences of testing. Writing that your questionnaire “has content validity” therefore reads as dated in an American methods section. Writing that content-related evidence supports the intended interpretation reads as current.
How reliability gets estimated: across time, items and raters
Reliability is estimated by repeating something and comparing the results. Test-retest reliability repeats the occasion, parallel-forms reliability repeats the instrument with an equivalent version, split-half and internal consistency repeat across the items of one questionnaire, and interrater reliability repeats across the people who assign the scores.
Which coefficient belongs to which approach is worth checking before you report a number. The Early Intervention Research Group at Northwestern University lists internal consistency with Cronbach’s alpha, test-retest with correlations, and interrater reliability with kappa or intraclass correlations. Reporting an alpha for a two-rater coding exercise is a common and avoidable mistake.
Quality criteria in qualitative research: the four criteria of Lincoln and Guba
Qualitative research in the English-speaking world is most often judged against the four criteria of Yvonna Lincoln and Egon Guba: credibility, transferability, dependability and confirmability, summarized as trustworthiness.
Presented in “Naturalistic Inquiry” (Sage, 1985), the four were designed as counterparts to the conventional criteria rather than as translations of them. Almost every English-language methods text still starts here, which is why an American thesis that cites a qualitative catalog usually cites this one.
| Criterion | The question behind it | Conventional counterpart |
|---|---|---|
| Credibility | Does the interpretation fit the material? | internal validity |
| Transferability | Can readers judge a transfer elsewhere? | external validity |
| Dependability | Did the process stay traceable? | reliability |
| Confirmability | Do the findings rest on the data? | objectivity |
The pairing in the right-hand column comes from Lincoln and Guba themselves, who introduce their four criteria as rough analogues of internal validity, external validity, reliability and objectivity. Practice follows from the wording. Transferability is served by thick description, so that readers can decide for themselves what carries over to their setting. Dependability and confirmability are served by an audit trail: decisions, memos and coding rules kept in a form a stranger could follow. Four years later, in “Fourth Generation Evaluation” (Sage, 1989), Guba and Lincoln added five further criteria under the heading authenticity, which are cited far less often.
💡 Tip
Ask a second person to code five of your interview passages independently and compare the outcome. Even that small sample shows whether your category system holds, and it gives you one defensible sentence for the methods chapter. Without such a comparison, confirmability stays an assertion.
Which other catalogs of qualitative quality criteria exist
Three further catalogs turn up regularly in student work: the eight big-tent criteria of Sarah J. Tracy, the six criteria of Philipp Mayring and the seven core criteria of Ines Steinke.
Choosing between them is not a matter of taste. Your method points to one of them, and so does the literature you cite. Working with qualitative content analysis makes Mayring the obvious partner, because method and criteria come from the same workshop. Judging a whole research process is where Steinke offers the broader frame. Writing for an American reader makes Tracy the most legible reference.
The eight big-tent criteria of Sarah J. Tracy
Sarah J. Tracy of Arizona State University proposes eight markers of excellent qualitative research: worthy topic, rich rigor, sincerity, credibility, resonance, significant contribution, ethics and meaningful coherence. Published in Qualitative Inquiry 16(10), pages 837 to 851, under the DOI 10.1177/1077800410383121, the model has become a common reference in American graduate programs.
What makes the model useful is a distinction the older catalogs blur. Tracy separates the end goals of good qualitative work from the means by which a researcher reaches them, so scholars from different paradigms can share one vocabulary without sharing one method. A compact summary with a table of means for each criterion is available in the encyclopedia entry by Tracy and Hinrichs.
Mayring and Steinke, the German-language catalogs you may meet
German-language literature works with two catalogs that have no direct English equivalent. Philipp Mayring names six general criteria of qualitatively oriented research: procedural documentation, argumentative validation of interpretations, adherence to rules, closeness to the subject, communicative validation and triangulation. They stand in chapter 6.3 of the seventh edition of 2023 (Beltz).
Ines Steinke sets out seven core criteria: intersubjective traceability, indication of the research process, empirical grounding, limitation, coherence, relevance and reflected subjectivity. The catalog is documented among others in the Dorsch dictionary of psychology (Hogrefe), which dates it to Steinke (1999). Knowing both pays off the moment your reading list includes German titles.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — free to start.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
Why there is no binding catalog for qualitative research
No binding catalog exists for qualitative research because the field has never agreed on one, and the disagreement is substantive rather than administrative. Qualitative designs are deliberately open, and a single fixed standard would restrict exactly the openness that gives them their purpose.
How unsettled the question remains is visible in the literature itself. Tracy and Hinrichs describe a methodological landscape holding a wide variety of quality concepts, one that even questions whether criteria are needed at all. Denzin (2008) coined the big-tent image that Tracy took up, and the image only works because several incompatible positions have to fit inside it.
American associations have answered with reporting standards instead of quality criteria. Levitt and colleagues published the Journal Article Reporting Standards for Qualitative Research in American Psychologist in 2018, the first time qualitative work entered APA Style at all. They prescribe what a manuscript has to disclose, not what counts as good research, which sidesteps the disagreement rather than settling it.
German-language sociology went the other way and kept proposing catalogs. Strübing, Hirschauer, Ayaß, Krähnke and Scheffer put forward five criteria in 2018 under the subtitle “Ein Diskussionsanstoß”, a starting point for discussion, and replies followed in the same journal. For your own work the lack of consensus is good news rather than bad.
Naming one catalog, justifying it and applying it consistently is worth more than listing all four.
Why German-language sources add a third criterion called Objektivität
German-language methodology counts a third main criterion alongside reliability and validity, called Objektivität, and it means independence of the result from the person who produced it.
Three forms are usually distinguished, and each of them names a stage of the study. Objectivity of administration means every respondent gets the same instruction under the same conditions. Objectivity of scoring means the rules for turning answers into values are fixed in advance. Objectivity of interpretation means the rules for reading those values are fixed as well. Lienert and Raatz set the trio out in 1998, and Moosbrugger and Kelava still treat the three forms as general criteria in their 2020 edition.

Nothing about that concern is missing from the American system, it simply lives somewhere else. Variability between qualified raters counts as a source of measurement error in the 2014 Standards and is therefore handled under reliability. Standardized administration and scoring are covered by a separate cluster of standards on test design, development, administration and scoring procedures. In the qualitative catalog of Lincoln and Guba the same worry surfaces once more, as confirmability.
Two practical consequences follow for your own writing. Finding no discussion of objectivity in an American source is not an omission by that source, it is a difference in how the field is organized. And when you render a German passage into English, do not translate Objektivität as objectivity, a word that English-language methods writing uses for a much broader epistemological stance. Describe what is meant instead: standardized administration, scoring and interpretation.
Typical mistakes with quality criteria in a thesis
Most points are lost in the methods chapter not through weak numbers but through claims that reach further than the data allow. Three mistakes account for the bulk of them.
What the three have in common is that a criterion gets asked to do a job it was never built for. Keeping the catalogs apart avoids all three almost by itself, and the methods chapter gets shorter in the process rather than longer.
Treating objectivity as a property of the room
Objectivity means independence from the researcher, not independence from the surroundings. A frequent claim is that a survey was objective because it took place in a neutral room. A neutral room is a condition of administration. What makes a study objective is that instruction, scoring rules and interpretation rules are fixed in advance and apply to everyone alike.
Calling validity a type instead of a body of evidence
Validity is not a property that a questionnaire either has or lacks. Sentences such as “the scale has construct validity” treat it as a possession and read as dated to an American reader. Validity applies to an interpretation of scores for a stated purpose, and it is supported by evidence, which is why the current standards speak of sources of validity evidence.
Listing a catalog without applying it to your own study
Naming the eight big-tent criteria is not yet an examination of your own quality. Graders look for whether you can say, for each criterion, how you met it and where you hit a limit. A sentence such as “member reflections were not possible within the timeframe, so interpretations rest on thick description and an audit trail” is worth more than a complete list without reference.
Conclusion
For quantitative research the question of quality criteria is settled, for qualitative research it is not, and the openness is the result of a live scholarly debate rather than a defect. Pick a catalog, name the source it comes from, and work it through your own data collection. That gives you the methods chapter you actually need.
Where to go next
- Need the three quantitative criteria in detail? Objectivity, Reliability, Validity: Definition and Example
- Unsure which approach fits your question? Qualitative vs. quantitative: definition and differences
- Running interviews instead of a survey? Qualitative Survey: Types, Interview Guide and Analysis
- Looking for the right analysis procedure? Data Analysis Methods: Choosing the Right One for Your Data
- Want the whole process in view? Fundamentals of Empirical Research
Our reading tip: Tracy, Sarah J. (2013): Qualitative Research Methods. Collecting Evidence, Crafting Analysis, Communicating Impact. Wiley-Blackwell: Hoboken. The book expands the big-tent model into a full course in qualitative craft and is the shortest serious way into the topic. For the quantitative side: American Educational Research Association, American Psychological Association and National Council on Measurement in Education (2014): Standards for Educational and Psychological Testing. Washington, DC.
Want every respondent to go through your survey the same way?
With empirio.ai, an online survey tool from Germany, everyone sees the same questionnaire in the same order, and the answers arrive already structured. Data is stored on servers in the European Union, and your institutional review board or research office decides what that means for your study.
