empirio.ai

Objectivity, Reliability, Validity: Definition and Examples

Your questionnaire can be perfectly consistent and still measure the wrong thing. Find out what objectivity, reliability and validity check and why the BPS model names only two.

Author at empirio.ai - Maria Malzewby Maria MalzewUpdated 25 August 2026Reading time 17 min

Picture two markers analysing the same completed questionnaires and arriving at two different results. The questionnaire is not badly written. Objectivity is simply missing.

Objectivity, reliability and validity are the three main quality criteria in quantitative research, and each of them checks something different. Objectivity means that the result does not depend on who runs, scores and interprets the study. Reliability means that an instrument measures accurately, keeping measurement error as small as possible. Validity means that a questionnaire captures the characteristic you actually care about, rather than a different one by accident. After reading, you can justify for your own study which criterion you meet and which you deliberately leave open.


📌 The key points at a glance

  • Objectivity means the result does not hang on the person analysing.
  • Reliability means the questionnaire measures with little measurement error.
  • Validity means the intended characteristic is what actually gets measured.
  • Reliability is a precondition for validity, not the other way round.
  • The BPS review model names validity, reliability and norms, not objectivity.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

Objectivity, reliability and validity: the three quality criteria at a glance

Objectivity, reliability and validity are the three main quality criteria used to judge how good a measuring instrument is. Taken in order, they answer three questions: is the measurement independent of the person running it, is it precise, and does it measure the right thing?

The order is not a matter of habit, because every criterion depends on the one before it. A study in which each marker arrives at a different value cannot be reliable in the first place, and an unreliable measurement can never be valid. The connection is stricter than most summaries suggest, and it deserves a chapter of its own.

Quality criterionThe question behind itWhat you pin it on
ObjectivityDoes the result depend on the person?fixed rules for administration, scoring, interpretation
ReliabilityDoes a repeat produce the same value?reliability coefficient between 0 and 1
ValidityIs the intended characteristic measured?content, construct and criterion related evidence

What most guides leave out: the triad is neither the whole list nor the same list everywhere. In the German speaking tradition, Helfried Moosbrugger and Augustin Kelava set out ten quality criteria in their textbook “Testtheorie und Fragebogenkonstruktion”, and alongside the familiar three they list scaling, norming, test economy, usefulness, reasonableness, resistance to faking and fairness, as you can read in chapter 2 of that textbook (Springer, 2012).

For an undergraduate dissertation, working with the three main criteria is normally enough. Just be aware that the restriction is a choice you are making. Whoever can at least name the further quality criteria in empirical research tends to sound noticeably more confident in the viva.

What the British Psychological Society actually assesses

British test standards do not run on a three part triad at all. Tests submitted to the Psychological Testing Centre of the British Psychological Society are, in the society’s own words, “reviewed independently by two Reviewers and two Editors against the European Federation of Psychologists Association Review Model for the Description and Evaluation of Psychological Tests”. Assessed there are the quality of the technical and user documentation, the quality of the test materials, and the validity, reliability and provision of norms of the test itself.

Two further details matter if you cite a published test in your dissertation. A BPS Registered Test carries a Certificate of Test Registration that stays valid for five years, so a registration is a dated statement rather than a permanent seal of approval. Submission is voluntary as well: publishers decide for themselves whether to have a test reviewed, and the absence of a BPS review therefore says nothing about how good a test is.

Why you meet “objectivity” in translated material but not in the BPS model

Objectivity does not appear as a criterion in the BPS and EFPA review model, and its work is done by two other things instead: standardisation, which fixes how a test is administered and scored, and the quality of the technical and user documentation, which is assessed in its own right. Much of the English language material on quality criteria is translated from German textbooks, which is why the three part triad keeps surfacing in English searches.

For your dissertation the practical consequence is small but worth knowing. Keep using objectivity as a heading where your module handbook and your supervisor expect it, since the underlying idea of removing the researcher from the result is uncontroversial. Only avoid claiming that objectivity is a criterion in the British review model, because a marker with a psychometrics background will notice.

Diagram of the three quality criteria objectivity, reliability and validity side by side

Objectivity: does the result depend on the person running the questionnaire?

Objectivity is given when a test measures the characteristic it measures independently of who administers and scores it, and when clear, user independent rules exist for interpreting the values. In practice that means the person running the study is left no room for manoeuvre.

In an online survey the criterion is easier to meet than in any oral procedure, because no conversation takes place between researcher and respondent. Standardised online surveys are popular in student projects for exactly that reason. The three aspects of objectivity go back to Gustav Lienert and Ulrich Raatz (1998), Moosbrugger and Kelava adopt them in that form, and a dissertation justifies each aspect separately.

Administration objectivity: the same conditions for every respondent

Administration objectivity is given when the result does not depend on who runs the study. The lever for it is standardisation: instructions, question order, timing and answer formats are fixed in advance and do not change between two participants. Moosbrugger and Kelava describe the ideal as a situation in which the respondent is the only source of variation. A survey distributed as a link and looking identical for everyone meets that almost automatically.

Scoring objectivity: the same answers, the same analysis

Scoring objectivity is given when different people arrive at the same result from identical answers. With closed questions and fixed response options nothing has to be interpreted, so the criterion causes no trouble. As soon as you work with open ended questions, scoring objectivity drops, because free text has to be sorted into categories. Coding free text is no reason to avoid it, but it does require a documented coding frame.

Interpretation objectivity: the same values, the same conclusion

Interpretation objectivity is given when different specialists draw the same conclusion from the same numerical value. Among the three aspects it is the weakest in practice, because norms or clear thresholds would have to exist for it. In a study of your own without a norming sample, none are available. What you do instead: fix the interpretation rules in advance and state them in the methods chapter, rather than aligning them with the result afterwards.

💡 Tip

Write down the coding rules for your open questions before you read the first answer. Whoever develops the coding frame on the finished data set can no longer rule out that the categories were chosen to fit the desired result. In a viva that is the most uncomfortable question of all.

Reliability: does a repeat measurement produce the same result?

Reliability describes the measurement precision of a test. A test is reliable when it measures the characteristic it measures without measurement error. The yardstick for it is the reliability coefficient, which lies between 0 and 1.

A coefficient of 1 means the measurement is free of measurement error, a coefficient of 0 that the value came about through measurement error alone. Moosbrugger and Kelava write that the reliability coefficient of a good test should not fall below 0.7. The figure circulates online without a source surprisingly often. It appears in chapter 2 of the Moosbrugger and Kelava textbook (2012) as a target figure, not as a cut off below which a scale would be unusable.

Four procedures are common in classical test theory for determining reliability. Which of them is realistic for your dissertation is decided less by statistics than by the question of whether you can reach your respondents a second time.

ProcedureHow it worksFeasible in a dissertation?
Test retest reliabilitythe same test at two points in time, values correlatedrarely, respondents are seldom reachable again
Parallel test reliabilitytwo equivalent forms of the same testhardly ever, the item pool is too small
Split half reliabilitytest split into halves, halves correlatedyes, with the Spearman Brown correction
Internal consistencyevery item as its own part test, usually Cronbach’s alphayes, the normal case

Parallel test reliability counts as the royal road in test theory, because it rules out practice and memory effects. For a bachelor’s or master’s dissertation it is nonetheless almost never feasible, since it demands two questionnaire versions leading to the same true values. In practice nearly everyone ends up with internal consistency.

Cronbach’s alpha: what the value delivers and what it does not

Cronbach’s alpha measures how strongly the items of a scale hang together, which makes it a measure of internal consistency. Every statistics package outputs the value, and partly for that reason it gets over interpreted. A high alpha does not prove that a scale measures only one characteristic, and it rises simply because a scale gains more items.

The assumption behind alpha is strict as well, because all items are supposed to load equally strongly on the same characteristic. Where that does not hold, alpha underestimates the true reliability. Klaas Sijtsma set out these limitations at length in the journal Psychometrika in 2009, and McDonald’s omega has often been named as the alternative since.

For a dissertation none of that means you have to compute omega. Rather, it means you place alpha in context with half a sentence instead of presenting it as proof. Precisely such qualifications separate a good methods discussion from an average one.

Diagram of the quality criterion reliability as a repeated measurement with the same result

Validity: does the questionnaire really measure the intended characteristic?

Validity is given when a test really measures the characteristic it is supposed to measure, and not some other one. Moosbrugger and Kelava call validity the most important quality criterion for test practice, because it decides how much any result can say.

Unlike reliability, validity is not shown by a single number. Instead it is assembled from several kinds of evidence, and depending on the application only some of them can be examined at all. The textbook distinguishes four aspects, and a methods chapter should call them by those names.

  • Content validity: do the items cover the characteristic representatively? The question is not calculated but answered through subject knowledge and expert judgement.
  • Face validity: does the claim of a test look justified to a layperson at first glance? Acceptance is at stake here, not explanatory power.
  • Construct validity: is the inference from answer behaviour to the underlying characteristic theoretically founded? Evidence comes from correlations with similar and with deliberately dissimilar instruments.
  • Criterion validity: can the test score predict behaviour outside the test situation? With a criterion available at the same time the term is concurrent validity, with one lying in the future, predictive validity.

The most common mistake: content validity and face validity are not the same

Frequently you read that content validity and face validity are two names for one thing. The claim is wrong, and the mix up is common enough that Moosbrugger and Kelava flag it in the textbook as a recognised risk of confusion. The difference lies in who judges, and on what basis. Content validity is a specialist judgement about whether the items cover the domain. Face validity is a layperson’s impression that a test looks plausible.

⚠️ Warning

Do not write “content validity, or face validity” in your dissertation. Whoever equates the two terms hands an open flank to anyone examining the methods chapter. Name content validity as the thing you justify on subject grounds, and face validity at most as an argument for acceptance among respondents.

A second blurred distinction concerns internal and external validity. Both terms come from the evaluation of study designs and answer whether a causal inference holds within a study and whether the findings generalise beyond it. Neither one is a subtype of content or construct validity, even though many summaries file them that way. Mixing the two systems costs exactly the precision that markers look for in a methods chapter.

Diagram of the quality criterion validity as a hit on the intended characteristic

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

How objectivity, reliability and validity hang together

Reliability is the precondition for validity, and the relationship runs in one direction only. Moosbrugger and Kelava put it as follows: objectivity and reliability supply only the favourable preconditions for high validity, and a test with low reliability cannot have high validity.

The quoted sentence from the textbook looks unremarkable and still carries the whole system. A reliable measurement can miss the point entirely, whereas an unreliable measurement can never be valid. Objectivity comes before both in practice, because a study that leaves wide scope in the scoring rarely produces stable values. The textbook does not state that as a formal condition, unlike the relationship between reliability and validity.

A bathroom scale that consistently reads two kilograms too high makes the difference concrete. Such a scale is objective, because it works by the same rules for every person. Such a scale is also reliable, because three measurements in a row produce the same figure. What it is not is valid, since it measures systematically wrong. Precisely that case, high reliability with low validity, is the most common one in questionnaires, and the coefficients do not reveal it.

What follows for the order in your methods chapter

  • Clarify first whether the study runs identically for every respondent.
  • Check next whether your scales are internally consistent.
  • Justify only last that the items cover the intended characteristic.
  • Report a low coefficient instead of leaving it out.

A reliable questionnaire measures dependably. Whether it measures the right thing is a separate question.

Do the three quality criteria also apply in qualitative research?

Objectivity, reliability and validity were developed for standardised measurement and cannot be transferred to qualitative procedures unchanged. A semi structured interview is not supposed to produce the same thing on every repetition, precisely because it responds to the individual person.

Qualitative social research has therefore developed criteria of its own. The best known proposal comes from Yvonna Lincoln and Egon Guba in “Naturalistic Inquiry” (Sage, 1985), with the four criteria credibility, transferability, dependability and confirmability. German language methods literature additionally points to the core criteria of Ines Steinke, which put the intersubjective traceability of the procedure at the centre.

For your own dissertation the practical consequence is manageable. Where you collect data qualitatively, do not write about Cronbach’s alpha, but document selection, procedure and analysis precisely enough for someone else to retrace the path. Which criteria sit side by side in the two research logics is set out in the comparison of quantitative and qualitative quality criteria.

Quality criteria in your dissertation: where they belong in the structure

Quality criteria have no fixed place in the structure of a dissertation. Two arrangements are common, and both are defensible: a subsection of its own in the chapter on research design, or a section in the reflection on your own procedure.

Which of the two fits depends on how much your choice of method needs justifying for your research question. Where the method stands at the centre, the criteria belong near the front in the design. Where the substantive findings matter more, they sit more naturally in the reflection. Check the point with your supervisor, since some departments state a fixed expectation in the module handbook or in your university’s regulations.

Not every dissertation collects its own data. Where you analyse an existing British survey instead, the Learning Hub of the UK Data Service, funded by the Economic and Social Research Council, explains how to handle weighting and complex sample designs. Reliability and validity of the instrument are not covered there, so the quality criteria for the questionnaire itself still have to come from the documentation of the original study.

Option 1: quality criteria as a subsection of the research design

Option 1 suits you where you justify your choice of method at length and the quality criteria form part of that justification. The advantage is that readers have settled the quality question before they see a single finding. The price is that you do not yet know at that point which scales will turn out weak, so expect to add a sentence in the reflection later on.

  1. Introduction with the research question
  2. Theory and state of research
  3. Research design, containing hypotheses, choice of method, population and sample, variables, question catalogue, and quality criteria with one subsection each for objectivity, reliability and validity
  4. Data collection
  5. Data analysis
  6. Findings
  7. Reflection on the procedure
  8. Conclusion

Option 2: quality criteria as a section in the reflection on your procedure

Option 2 suits you where the limits of your study only become describable in the light of the findings. The structure stays the same, except that the quality criteria move into point 7 and get discussed there together with the limitations of the work. The advantage is honesty, because you write about coefficients you actually obtained rather than about intentions. The price is that the methods critique sits a long way back and is easily skimmed over.

In both arrangements the same standard applies: what gets marked is not whether a study meets all three criteria, but whether it becomes clear which are met and where the limits lie.

Overview of the points in a dissertation where objectivity, reliability and validity appear

Typical mistakes with the quality criteria

Most marks lost in the methods chapter come not from weak coefficients but from statements claiming more than the data supports. Three mistakes turn up particularly often.

All three mistakes share one cause: a quality criterion gets claimed for something it was never built for. Whoever keeps the three criteria properly apart avoids the mistakes almost automatically, and the methods chapter gets shorter rather than longer in the process.

Mistake 1: selling high reliability as proof of validity

A Cronbach’s alpha of 0.89 says that the items of a scale hang together closely. Nothing follows from the value about whether those items capture the intended characteristic. Whoever writes after the alpha value that the scale is therefore valid draws exactly the inference test theory rules out. The way out is an additional paragraph on content validity, justifying why the items cover the domain.

Mistake 2: treating objectivity as settled because the survey runs online

An online survey largely secures administration objectivity, because every respondent sees the same thing. Scoring and interpretation objectivity remain untouched by it. As soon as open questions appear in the questionnaire or thresholds get chosen freely, room for manoeuvre returns. The way out is to name all three aspects separately, rather than ticking objectivity off wholesale.

Mistake 3: keeping quiet about weak values

A reliability coefficient of 0.58 is no reason to drop a scale from the dissertation or to leave the value unreported. What gets marked is the engagement with the value, not the value itself. Whoever names it, explains what probably caused it and phrases the interpretation more cautiously shows exactly the methodological maturity being looked for. Values kept quiet, by contrast, surface at the latest in the viva.

Conclusion

Objectivity, reliability and validity are not hurdles you have to clear, but three questions you answer. Full compliance is the exception even in professional studies, and nobody expects it in a dissertation.

What is expected is that you can keep the three criteria apart and say openly where your study reaches its limits. Knowing that British test standards run on validity, reliability and norms rather than on the triad simply gives you one more precise sentence for the viva.

What to do next

Our reading tip: Moosbrugger, Helfried and Kelava, Augustin (eds.) (2012): Testtheorie und Fragebogenkonstruktion. 2nd edition. Springer: Berlin and Heidelberg. Chapter 2 covers all ten quality criteria and is freely available as a sample chapter. For the British side, the Psychological Testing Centre of the British Psychological Society sets out what a reviewed test has to document.


Want your survey to run to the same standard from the very first answer?

With empirio.ai, an online survey tool from Germany, every participant sees the same questionnaire in the same order, which largely secures administration objectivity. Scoring and interpretation objectivity stay your job, because both hang on your rules rather than on the tool.

Create a survey for free

Frequently asked questions

Reliability is the precision of a measurement, validity its soundness. A questionnaire is reliable when a repeat produces the same value, and valid when it captures the intended characteristic. Both are not the same thing: a bathroom scale that consistently reads two kilograms too high is reliable, but it is not valid.

Objectivity is the precondition for reliability, and reliability the precondition for validity. Moosbrugger and Kelava state that a test with low reliability cannot have high validity. The reverse does not hold, because a dependable measurement can still miss the characteristic it was meant to capture entirely.

Moosbrugger and Kelava name 0.7 as the value the reliability coefficient of a good test should not fall below. The figure is a rule of thumb rather than a fixed threshold. In a dissertation what counts is less the size of the value than whether it is reported and placed in context.

Objectivity is divided into administration, scoring and interpretation objectivity. Administration objectivity means every respondent meets the same conditions. Scoring objectivity means identical answers are analysed identically by different people. Interpretation objectivity means the same conclusion is drawn from the same numerical value. A dissertation justifies all three separately.

No, content validity and face validity are two different aspects, and Moosbrugger and Kelava name the confusion explicitly in their textbook. Content validity is a specialist judgement about whether the items cover the domain representatively. Face validity is a layperson impression that a test looks plausible, which concerns acceptance rather than explanatory power.

The British Psychological Society has tests reviewed against the EFPA Review Model, covering the quality of the technical and user documentation, the quality of the test materials, and the validity, reliability and provision of norms of the test. Objectivity is not a separate criterion there, and a BPS Registered Test stays registered for five years.

Quantitative quality criteria cannot be transferred to qualitative procedures unchanged, because a semi structured interview is meant to run differently each time. Lincoln and Guba propose credibility, transferability, dependability and confirmability in Naturalistic Inquiry (1985). German language literature additionally refers to the core criteria of Ines Steinke, centred on intersubjective traceability.

Two placements are common: a subsection of its own in the chapter on research design, or a section in the reflection on your own procedure. Where the choice of method stands at the centre of the work, the criteria belong near the front. Where substantive findings matter more, the reflection is the more natural home.

You might also be interested in

Empirical Research

Fundamentals of Empirical Research | empirio

Empirical does not mean numbers, and an existing dataset is not a shortcut. Find out which route fits your research question and how to justify it in your methods chapter.

Glossary

What does validity mean? Definition and example

Validity (= correctness or accuracy of the measurement) reflects the extent to which a research work in its entirety actually achieves the results that correspond to the stated research objective.