Research is only worth something if its findings are sound rather than misleading. That is why measurement accuracy, better known as reliability, is built into an empirical study right from the planning stage, so that errors creep in as rarely as possible.
Create a survey for free
With empirio.ai you can create a modern online survey in minutes — with hosting in the EU.
- AI-built survey
- Adjust by drag & drop
- Real-time analysis
General definition of reliability
Reliability (the dependability or consistency of a measurement) is a key quality criterion of any piece of research and describes how precisely a study measures what it sets out to measure. If the research were repeated at a different time under the same conditions, the researchers should arrive at the same or at least comparable results.
A study has high reliability when its findings are as free as possible from random error. If the measurements cannot be reproduced, the research is not considered dependable, in other words its reliability is low.
Example of reliability
A digital scale is a dependable measuring instrument with high reliability, because it shows the same body weight every time, even when you step on it again and again. A researcher estimating participants' weight by eye during an observation would be far less consistent, and that estimate would be rated as having low reliability.
Testing reliability in research
In research practice, reliability can be estimated with several methods. The most common are:
Test-retest method
The same measurement is repeated so that the two rounds can be compared. If the same instrument produces the same results with the same participants on both occasions, it is considered reliable. Because the method takes a lot of effort, it only really pays off in larger studies.
Parallel-forms method
Two equivalent versions of a measuring instrument, for example two questionnaires that cover the same construct with different but comparable questions, are used for the same research question at the same time. If both versions produce comparable results, the instrument is considered reliable. The catch is that it is not easy to develop two versions that really are equivalent.
Split-half method
A variant of the parallel-forms idea: instead of building two separate versions, you split a single instrument, such as a questionnaire, into two comparable halves after it has been completed, for instance the odd-numbered and the even-numbered items. If both halves lead to similar results, the instrument is considered reliable. Because it needs only one round of data collection, this method is very popular.
With all of these methods, the degree of reliability is calculated and expressed as a correlation coefficient: the closer it is to 1, the more reliable the measuring instrument.
In smaller studies, especially in student projects such as an undergraduate dissertation, it is not always possible to test reliability with the methods above. As long as you have chosen a suitable research method (your measuring instrument) and applied it with care, that level of reliability is usually sufficient.