empirio.ai

What Does Reliability Mean? Definition and Example

Reliability (= dependability or trustworthiness of measurement) is a crucial quality characteristic of research work and refers to the accuracy of the investigation carried out.

by Maria MalzewUpdated 25 August 2023Reading time 2 min

Research is only worth something if its findings are sound rather than misleading. That is why measurement accuracy, better known as reliability, is built into an empirical study right from the planning stage, so that errors creep in as rarely as possible.

Create a survey for free

With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.

  • AI-built survey
  • Adjust by drag & drop
  • Real-time analysis
Start for free

General definition of reliability

Reliability (the dependability or consistency of a measurement) is a key quality criterion of any piece of research and describes how precisely a study measures what it sets out to measure. If the research were repeated at a different time under the same conditions, the researchers should arrive at the same or at least comparable results.

A study has high reliability when its findings are as free as possible from random error. If the measurements cannot be reproduced, the research is not considered dependable, in other words its reliability is low.

Example of reliability

A digital scale is a dependable measuring instrument with high reliability, because it shows the same body weight every time, even when you step on it again and again. A researcher estimating participants' weight by eye during an observation would be far less consistent, and that estimate would be rated as having low reliability.

Diagram illustrating reliability as a quality criterion in quantitative research

Testing reliability in research

In research practice, reliability can be estimated with several methods. The most common are:

  1. Test-retest method
  2. Parallel-forms method
  3. Split-half method

Test-retest method

The same measurement is repeated so that the two rounds can be compared. If the same instrument produces the same results with the same participants on both occasions, it is considered reliable. Because the method takes a lot of effort, it only really pays off in larger studies.

Parallel-forms method

Two equivalent versions of a measuring instrument, for example two questionnaires that cover the same construct with different but comparable questions, are used for the same research question at the same time. If both versions produce comparable results, the instrument is considered reliable. The catch is that it is not easy to develop two versions that really are equivalent.

Split-half method

A variant of the parallel-forms idea: instead of building two separate versions, you split a single instrument, such as a questionnaire, into two comparable halves after it has been completed, for instance the odd-numbered and the even-numbered items. If both halves lead to similar results, the instrument is considered reliable. Because it needs only one round of data collection, this method is very popular.

With all of these methods, the degree of reliability is calculated and expressed as a correlation coefficient: the closer it is to 1, the more reliable the measuring instrument.

In smaller studies, especially in student projects such as an undergraduate dissertation, it is not always possible to test reliability with the methods above. As long as you have chosen a suitable research method (your measuring instrument) and applied it with care, that level of reliability is usually sufficient.

You might also be interested in

Glossary

What does validity mean? Definition and example

Validity (= correctness or accuracy of the measurement) reflects the extent to which a research work in its entirety actually achieves the results that correspond to the stated research objective.

Glossary

Empirical: Meaning, Definition and Examples

Empirical sounds like statistics and large samples. In fact the word began life as an insult, and what makes a dissertation empirical is where its data came from.