Reliability & Validity

Introduction to Reliability & Validity
🎯 The Dual Pillars of Research Rigor
Psychological research explores cause-and-effect relationships and constructs theoretical frameworks. Because psychological constructs (e.g. anxiety, intelligence, motivation) are abstract, establishing Reliability (Consistency) and Validity (Accuracy) is crucial for trustworthy conclusions.
If a thermometer reads 98.6°F three consecutive times for the same healthy person under identical conditions, the thermometer is reliable.
If a weighing scale shows 70 kg repeatedly for an object, it is consistent. But if it reads 70 kg for an object that is actually 60 kg, it is reliable but invalid!
Does the test yield repeatable, identical results under identical measurement conditions?
Does the test genuinely measure what it claims to measure, yielding meaningful inferences?
Meaning, Definitions & Measurement Error
📚 Classic Definitions of Reliability
📐 Classical Test Theory: $X = T + E$
Every psychological measurement involves some degree of error. In psychometrics, error reflects unavoidable measurement inaccuracies rather than human mistakes. The goal of research design is to minimize $E$ to maximize true score variance.
Methods of Estimating Reliability: External Consistency
External Consistency Procedures compare results obtained from two independent data collection processes across time or test forms.
Test-Retest Reliability
The same test is administered twice to the same group of participants after a specified time interval. The Pearson correlation ($r$) between the two sets of scores indicates temporal stability.
- Memory Effect: Participants recall past answers.
- Practice Effect: Scores improve due to familiarity.
- Participant Attrition: Missing subjects on re-test.
Parallel Forms Reliability
Two equivalent versions (Form A & Form B) of a test are administered to the same group. The correlation between scores on both forms estimates equivalence and consistency.
- Development Difficulty: Creating two truly matched forms with identical means/variances is extremely labor-intensive.
- Often replaced by internal consistency procedures.
Methods of Estimating Reliability: Internal Consistency
Internal Consistency Procedures evaluate how well items within a single test measure the same underlying construct on a single testing session.
Split-Half Reliability
The test is divided into two halves (e.g. odd vs. even items) and half-scores are correlated. Because shortening tests reduces reliability, the Spearman-Brown Prophecy Formula is applied to estimate full-length reliability:
Kuder-Richardson (KR-20)
Used for tests containing dichotomous items (e.g., True/False, Correct/Incorrect, 0/1). It evaluates inter-item consistency across all items simultaneously without requiring test splitting.
Cronbach's Alpha ($\alpha$)
The gold standard for multi-choice or Likert-scaled items. Measures average covariance among item pairs. Values range from $0.00$ to $1.00$ ($\ge 0.70$ acceptable, $\ge 0.80$ good).
Comparison of Reliability Estimators
| Reliability Estimator | Primary Use Case | Key Advantages | Key Limitations |
|---|---|---|---|
| Inter-Rater Reliability | Observational studies with multiple observers | Ensures consistency across observers | Requires multiple observers; resource-heavy |
| Parallel Forms Reliability | Tests with two equivalent versions | Reduces memory & practice effects | Time-consuming to build parallel forms |
| Cronbach's Alpha ($\alpha$) | Multi-item Likert scales / questionnaires | Handles multi-response continuous items | Assumes all items measure same construct |
| Test-Retest Reliability | Experimental & longitudinal designs | Demonstrates score stability over time | Affected by memory, practice, & time changes |
Meaning & Definitions of Validity
🎯 What is Validity?
Validity is the degree to which a test measures what it claims to measure. It ensures that the test results are meaningful, appropriate, and scientifically useful for drawing inferences.
📚 Classic Definitions of Validity
The correlation coefficient between test scores and an independent criterion measure.
External standard or behavior used to validate whether the test yields accurate real-world predictions.
6 Core Types of Validity in Research
Content Validity
Assesses whether test items cover the entire domain of construct (e.g., spelling test covering 3rd grade syllabus). Involves expert panel evaluation.
Criterion-Related Validity
Evaluates correlation with external criteria. Includes Concurrent Validity (same-time criteria) and Predictive Validity (future outcomes).
Construct Validity
Ensures test measures theoretical construct. Requires Convergent Validity (correlates with related constructs) and Discriminant Validity (does not correlate with unrelated constructs).
Face Validity
Superficial judgment of whether a test appears valid on the surface. Based on subjective impression.
Internal Validity
Degree to which changes in Dependent Variable are genuinely caused by Independent Variable (causal inferences).
External Validity
Extent to which findings can be generalized across other settings, populations, and time periods.
Threats to Internal & External Validity
🎯 7 Threats to Internal Validity
- Confounding Variables: Uncontrolled factors affecting DV.
- Selection Bias: Pre-existing group differences.
- History & Maturation: External events or natural aging/growth.
- Repeated Testing & Instrument Decay: Practice effects or tool drift.
- Mortality (Attrition): Subject dropout during study.
- Experimenter Bias: Unintentional researcher influence.
🌍 5 Threats to External Validity
- Aptitude-Treatment Interaction: Unique sample traits non-representative of population.
- Situational Factors: Environmental factors restricting generalisation.
- Pre-Test Effects: Sensitizing subjects to intervention.
- Post-Test Effects: Findings tied strictly to specific post-test conditions.
- Rosenthal Effects: Experimenter expectancy limiting generalizability.
Keep Listening & Keep Learning!

