empirio.ai

Reliability and Validity: Quality Criteria in Research

Quality criteria for qualitative research come from a different tradition than the quantitative ones. Find out which catalogue suits your method, from Lincoln and Guba to Yardley.

Author at empirio.ai - Maria Malzewby Maria MalzewUpdated 31 August 2026Reading time 12 min

Two students analyse the same twelve interviews and reach two different conclusions. Deciding which of the two analyses is the better one is the job of quality criteria, and qualitative research uses a different set from quantitative research.

Quality criteria are the standards by which a piece of research is judged usable. In the English-language tradition the pair is reliability and validity, and the “Standards for Educational and Psychological Testing” add fairness as a third foundation. Qualitative research has no single agreed catalogue. After this guide you can name the criteria your own study is judged on and show where it reaches its limits.


📌 The key points

  • Reliability and validity carry the quantitative tradition in English-language methodology.
  • AERA, APA and NCME name validity, reliability and fairness as foundations.
  • Lincoln and Guba propose four trustworthiness criteria for qualitative work.
  • Yardley (2000) offers four flexible principles rather than a checklist.
  • German-language sources add a third criterion, Objektivität, with no English counterpart.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

What are quality criteria in research?

Quality criteria are the standards that let you judge whether a piece of research is scientifically usable. They answer one question: does the result come from the procedure that was used, or from the person who used it?

The difference between an everyday observation and a scientific finding sits exactly here. “The mood in the team has turned” is a perception. Only once it becomes clear how that statement came about, with which instrument, from which people and under which rules, does it turn into a finding that colleagues can argue about on professional grounds.

A quality criterion is a scholarly convention rather than a rule. No law and no examination regulation prescribes which criteria a dissertation has to test. Markers do not check whether you meet all of them, but whether you can name the ones you meet and the points where your data reach their limit. Your department or supervisor has the final say.

Which criteria apply at all depends on the research tradition you are writing in. English-language methodology builds on reliability and validity, German-language textbooks add a third criterion, and qualitative research replaces the whole set with catalogues of its own. Where the criteria belong in your structure is covered in the guide to Objectivity, Reliability, Validity: Definition and Examples.

Quality criteria in quantitative research: reliability and validity

Quantitative research is judged on reliability and validity, and the current professional standards name fairness alongside them as a third foundation. Reliability asks how much measurement error a score carries, validity asks whether the intended attribute is being measured at all.

CriterionThe question behind itWhat you show it with
ValidityIs the intended attribute measured?evidence from content, structure and related variables
ReliabilityHow large is the measurement error?a reliability coefficient between 0 and 1
FairnessDoes the procedure disadvantage a group?evidence across the relevant subgroups

Reliability is a precondition for validity, and the reverse does not hold. A set of scales that is consistently two kilograms out measures reliably and is still wrong. How each criterion is tested in practice is set out in the guide to Objectivity, Reliability, Validity: Definition and Examples.

The reference work behind that ordering is the “Standards for Educational and Psychological Testing”, published jointly in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education. Part I of the volume, headed Foundations, runs to three chapters: Validity, Reliability/Precision and Errors of Measurement, and Fairness in Testing.

Sources of validity evidence, and why the old three-way split survives

Content, construct and criterion validity are the three labels most textbooks still teach. The 2014 Standards deliberately dropped them, describing validity as a unitary concept and listing five sources of evidence instead: test content, response processes, internal structure, relations to other variables, and consequences of testing. Both vocabularies are in circulation, so name the source your terms come from.

Ways of estimating reliability

Four estimation methods dominate the textbooks. Test-retest correlates the same instrument at two points in time, parallel-forms compares two equivalent versions, split-half correlates two halves of one instrument, and internal consistency, usually reported as Cronbach’s alpha, asks how far the items of a scale move together. Each answers a slightly different question, so state which one you used.

The four trustworthiness criteria of Lincoln and Guba

Qualitative research in the English-speaking world is most often judged against four criteria: credibility, transferability, dependability and confirmability. Yvonna Lincoln and Egon Guba set them out in “Naturalistic Inquiry” (Sage, 1985) and grouped them under the heading trustworthiness.

Lincoln and Guba did not set out to translate the quantitative criteria but to replace them. A guided interview is not supposed to produce the same answers on repetition, because it responds to the individual in front of the researcher. What can be asked instead is whether the interpretation fits the material and whether a reader can follow how it was reached.

CriterionWhat it asksQuantitative counterpart
CredibilityDoes the interpretation fit the material?internal validity
TransferabilityDoes the finding travel to other settings?external validity
DependabilityIs the process documented and stable?reliability
ConfirmabilityDo the findings rest on the data?objectivity

The pairing in the right-hand column comes from Lincoln and Guba themselves, who introduce their four criteria as rough analogues of internal validity, external validity, reliability and objectivity. Reading that column as a translation aid rather than an equation keeps you out of trouble. Four years later the same authors added five further criteria under the heading authenticity in “Fourth Generation Evaluation” (Sage, 1989), which student work takes up far less often.

💡 Tip

Ask a second person to code five of your interview passages independently and compare the result. Even a sample that small shows whether your category system holds, and it gives you a defensible sentence for the methods chapter. Without such a check, dependability stays a claim rather than a demonstration.

Which other catalogues of qualitative quality criteria exist

Three further catalogues turn up regularly in dissertations: the four principles of Lucy Yardley, the six criteria of Philipp Mayring and the seven core criteria of Ines Steinke. Choosing between them is a question of method rather than taste.

Working with interpretative phenomenological analysis or health psychology, you will meet Yardley almost immediately. Citing German-language literature, you will meet Mayring and Steinke instead, and their vocabulary does not map onto the English terms one to one. Naming the catalogue you follow, and the source it comes from, matters more than picking the fashionable one.

The four principles of Lucy Yardley

Lucy Yardley, then at the University of Southampton, proposed four open principles in “Dilemmas in qualitative health research” (Psychology & Health, 2000): sensitivity to context, commitment and rigour, transparency and coherence, and impact and importance. Her point was that quality markers should prompt reflection rather than work as a checklist that constrains the design.

A recent open-access review written by seven authors at UK universities places those principles alongside its own general criteria and quotes them in full. Elida Cena and colleagues published “Quality Criteria: General and Specific Guidelines for Qualitative Approaches in Psychology Research” in the International Journal of Qualitative Methods in 2024.

Mayring and Steinke, the catalogues behind German-language sources

Philipp Mayring names six general quality criteria for qualitative research: documentation of procedure, argumentative grounding of interpretations, adherence to rules, closeness to the object of study, communicative validation and triangulation. They appear in “Einführung in die qualitative Sozialforschung”, most recently in the seventh edition of 2023 (Beltz), in chapter 6.3.

Ines Steinke sets out seven core criteria: intersubjective comprehensibility, indication of the research process, empirical grounding, limitation, coherence, relevance and reflected subjectivity. The Dorsch Lexikon der Psychologie (Hogrefe) dates the catalogue to Steinke (1999). Two years circulate for it, because the monograph of 1999 was followed shortly afterwards by a handbook chapter.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

Why qualitative research has no binding catalogue

No binding catalogue exists because the field has never agreed on one, and the reason lies in the subject matter. Qualitative methods are deliberately open, and a single fixed standard would restrict the openness that gives them their purpose.

How live the argument still is shows in the most recent attempt. Jörg Strübing, Stefan Hirschauer, Ruth Ayaß, Uwe Krähnke and Thomas Scheffer proposed five criteria in the Zeitschrift für Soziologie in 2018: appropriateness to the object, empirical saturation, theoretical penetration, originality and textual performance. Their paper carries the subtitle “Ein Diskussionsanstoß”, a prompt for discussion, and replies followed in the same journal.

Nothing in that debate stops you finishing your methods chapter. Pick the catalogue that matches your method, cite the source it comes from, and work through it point by point on your own data. Markers reward a justified choice far more than an exhaustive list.

Naming one catalogue, justifying it and applying it consistently is cleaner work than listing all four.

Why German-language sources add a third criterion called Objektivität

German-language methodology names a third main criterion, Objektivität, that the English-language standards do not list at all. Meeting it in a German source is not a sign that the English literature has overlooked something, but a difference in where the same concern is filed.

Whether you need the term depends on the literature you cite. Writing in English about a German-language framework, translate it with care, because “objectivity” in ordinary English carries a philosophical meaning that the technical term does not.

What Objektivität means in the German system

Objektivität means independence from the researcher, not independence from the setting. German textbooks split it three ways: standardised administration, so that every respondent receives the same instructions; standardised scoring, so that the same answers produce the same values; and standardised interpretation, so that the same values lead to the same conclusion. A quiet room does not make a study objective.

Where the same concern sits in English-language methodology

English-language methodology handles the same problem inside reliability and inside the requirements for a standardised testing procedure. The 2014 Standards define reliability and precision as the consistency of scores across instances of the testing procedure, and they make the estimate depend on which sources of variation are allowed, raters among them. Agreement between raters is therefore a reliability question rather than a criterion of its own.

Teaching material shows the same split. The Open University sets out research quality in its free course on classroom research under exactly two headings, reliability and validity, and suggests reading reliability in naturalistic settings as dependability. On the qualitative side, confirmability in Lincoln and Guba covers the same ground: findings should rest on the data rather than on the researcher’s preferences.

Typical mistakes with quality criteria in a dissertation

Most marks are lost in the methods chapter not through weak coefficients but through claims that go further than the data allow. Three mistakes come up again and again.

Common to all three is that a criterion is asked to do a job it was never built for. Keeping the catalogues apart avoids all three almost automatically, and the methods chapter gets shorter rather than longer.

Reading a reliable measure as a valid one

A high reliability coefficient says that your instrument measures consistently, and nothing more. Consistency is a precondition for validity, never a substitute for it. An instrument can produce the same score every time and still capture a different attribute from the one your research question names, which is why validity needs evidence of its own.

Citing a catalogue with the wrong year or the wrong authors

Lincoln and Guba published the four trustworthiness criteria in 1985, while the five authenticity criteria appeared under Guba and Lincoln in 1989. Mayring’s book has run to seven editions since 1990, so several years circulate for the same six criteria. Cite the edition you actually held, and keep one year per source in your bibliography.

Listing a catalogue without applying it to your own study

Listing Yardley’s four principles is not yet an assessment of your own quality. Markers look for what you can say about each principle: how you met it and where you ran into a limit. A sentence such as “respondent validation was not possible within the timeframe, so interpretations were grounded argumentatively” is worth more than a complete list with no link to the study.

Conclusion

Quantitative research has a settled answer to the question of quality criteria, qualitative research does not, and the gap is the result of a live debate rather than a defect. Choose a catalogue, name the source it comes from, and work it through on your own data. Markers are looking for the reasoning, not for the length of the list.

Where to go next

Our reading tip: American Educational Research Association, American Psychological Association and National Council on Measurement in Education (2014): Standards for Educational and Psychological Testing. AERA: Washington DC. Part I covers validity, reliability and fairness and is the shortest authoritative starting point. For the qualitative side: Lincoln, Yvonna S. and Guba, Egon G. (1985): Naturalistic Inquiry. Sage: Newbury Park.


Want every respondent to work through the same questionnaire?

With empirio.ai, an online survey tool from Germany, all participants see the same questions in the same order, and the answers arrive already structured, which covers the administration side of standardisation. Responses are stored in the EU, which under UK GDPR is a shorter route than a chain of adequacy decisions and standard contractual clauses, although your organisation’s data protection officer decides the individual case. Scoring and interpretation rules stay with you.

Create a survey for free

Frequently asked questions

Quality criteria are the standards that show whether a piece of research is scientifically usable. Reliability and validity carry the quantitative tradition in English-language methodology, and the 2014 Standards for Educational and Psychological Testing name fairness alongside them. Qualitative research uses catalogues of its own, most often the four trustworthiness criteria of Lincoln and Guba.

Reliability describes how consistently an instrument measures, validity describes whether it measures the attribute you intended. Reliability is a precondition for validity, because a measure that varies at random cannot capture anything dependably. The reverse does not hold: a set of scales that is always two kilograms out measures reliably and is still wrong.

Qualitative research has no single agreed catalogue, so several competing proposals are in use. Lincoln and Guba name four trustworthiness criteria, Lucy Yardley four flexible principles, Philipp Mayring six criteria and Ines Steinke seven. For a dissertation it is enough to pick one catalogue, cite the source and apply it consistently to your own data.

Lincoln and Guba group credibility, transferability, dependability and confirmability under the heading trustworthiness. Credibility asks whether the interpretation fits the material, transferability whether the finding travels to another setting, dependability whether the process is documented and stable, and confirmability whether the findings rest on the data. The four appear in Naturalistic Inquiry, published by Sage in 1985.

Objektivität is a convention of German-language methodology rather than an international standard. Meaning independence from the researcher, it covers standardised administration, scoring and interpretation. The AERA, APA and NCME Standards of 2014 have no chapter on it, because the same concern sits inside reliability, in agreement between raters and in the requirements for a standardised procedure.

Lucy Yardley proposes sensitivity to context, commitment and rigour, transparency and coherence, and impact and importance. She published them in Dilemmas in qualitative health research in Psychology and Health in 2000. Yardley presents them as open prompts rather than a checklist, so that the flexibility qualitative designs depend on is not lost.

Reliability and validity cannot be carried over to qualitative research unchanged. A guided interview is not meant to produce the same answers on repetition, because it responds to the individual being interviewed. Lincoln and Guba therefore offer dependability in place of reliability and credibility in place of internal validity, which keeps the underlying question without the measurement assumptions.

You might also be interested in

Empirical Research

Fundamentals of Empirical Research | empirio

Empirical does not mean numbers, and an existing dataset is not a shortcut. Find out which route fits your research question and how to justify it in your methods chapter.