Reliability in statistics refers to how consistently a measurement instrument produces the same result when applied to the same thing. Two statistics dominate reliability analysis: Cronbach’s alpha for internal consistency and Cohen’s kappa for inter-rater agreement. They answer entirely different questions, use different formulas, and apply in different research situations, yet researchers sometimes conflate them. This article explains both from the ground up, with step-by-step numeric examples for each.

Cronbach’s alpha (also written Cronbach alpha in some software output) quantifies how well a group of survey or test items all measure the same underlying construct. Cohen’s kappa quantifies how well two independent raters agree on categorical classifications, after subtracting the agreement that would occur purely by chance. Learning when and how to apply each one is essential for designing and reporting defensible measurements in psychology, medicine, education, and the social sciences.


What Is Reliability in Statistics?

A measurement is reliable when it gives consistent results. A bathroom scale is reliable if it reads the same weight twice in a row under identical conditions. A psychological questionnaire is reliable if participants who genuinely share the same level of a trait score similarly on it. Reliability is a necessary but not sufficient condition for validity: a scale cannot accurately measure what it claims to measure if it cannot even measure anything consistently.

Researchers distinguish several types of reliability:

  • Test-retest reliability: does the instrument produce the same scores on two separate occasions for the same individuals?
  • Split-half reliability: if you randomly split a scale into two halves, do both halves produce similar scores?
  • Internal consistency reliability: do all items on a scale correlate with each other, indicating they tap a common construct? Cronbach’s alpha is the standard index for this.
  • Inter-rater reliability: when two or more observers independently classify the same cases, how much do they agree beyond what chance would predict? Cohen’s kappa is the standard index for this.

Identifying which type of reliability your study demands is the first analytical decision. The choice of statistic follows from that.


Internal Consistency: Cronbach’s Alpha

Internal consistency reliability answers the question: “Do all of these items hang together as a coherent scale?” Consider a five-item questionnaire designed to measure job satisfaction. If a respondent scores high on item 1 (“I enjoy my work”), they should tend to score high on item 2 (“I feel valued at my job”), because both items tap the same underlying construct. Cronbach’s alpha captures that coherence across the entire item set simultaneously.

Alpha ranges from 0 to 1. A value of 0 means the items share no common variance whatsoever; a value of 1 means every item is a perfect linear function of every other (total redundancy). In practice, well-designed scales typically fall between 0.70 and 0.95.

The Cronbach’s Alpha Formula

The standard formula, introduced by Lee Cronbach in 1951, is:

α = (k / (k − 1)) × (1 − ΣVᵢ / Vₜ)

Where:

  • k = the number of items on the scale
  • ΣVᵢ = the sum of each individual item’s variance
  • Vₜ = the variance of the composite total score (the sum across all items)

An equivalent covariance form reveals the underlying logic more clearly:

α = (k × c̄) / (v̄ + (k − 1) × c̄)

Where is the average inter-item covariance and is the average item variance. When items share substantial covariance relative to their individual variances, alpha is high. When items are largely unrelated to each other, alpha is low. Both forms of the formula are algebraically equivalent; most statistical software uses the variance form internally.

Step-by-Step Worked Example

Consider a three-item satisfaction scale administered to four employees, each item scored on a 1–5 Likert scale:

RespondentItem 1Item 2Item 3Total
A54514
B3339
C45413
D2226

Step 1 — Compute each item’s variance.

The Cronbach formula uses population variance (dividing by n, not n − 1), because both the item variances and the total variance use the same denominator and it cancels in the ratio.

Item 1 scores: 5, 3, 4, 2 → mean = 14/4 = 3.5

Deviations from mean: 1.5, −0.5, 0.5, −1.5 → squared: 2.25, 0.25, 0.25, 2.25 → sum = 5.00

V₁ = 5.00 / 4 = 1.25

Item 2 scores: 4, 3, 5, 2 → mean = 14/4 = 3.5

Deviations: 0.5, −0.5, 1.5, −1.5 → squared: 0.25, 0.25, 2.25, 2.25 → sum = 5.00

V₂ = 5.00 / 4 = 1.25

Item 3 scores: 5, 3, 4, 2 → mean = 14/4 = 3.5

Deviations: 1.5, −0.5, 0.5, −1.5 → squared: 2.25, 0.25, 0.25, 2.25 → sum = 5.00

V₃ = 5.00 / 4 = 1.25

Step 2 — Sum the item variances.

ΣVᵢ = 1.25 + 1.25 + 1.25 = 3.75

Step 3 — Compute the variance of the total score.

Total scores: 14, 9, 13, 6 → mean = 42/4 = 10.5

Deviations: 3.5, −1.5, 2.5, −4.5 → squared: 12.25, 2.25, 6.25, 20.25 → sum = 41.00

Vₜ = 41.00 / 4 = 10.25

Step 4 — Apply the Cronbach’s alpha formula.

α = (3 / (3 − 1)) × (1 − 3.75 / 10.25)
α = 1.5 × (1 − 0.366)
α = 1.5 × 0.634
α ≈ 0.951

A Cronbach’s alpha of 0.95 signals excellent internal consistency. The three items are measuring the same construct in a highly coherent way.


Interpreting Cronbach’s Alpha

The most widely used interpretation benchmarks come from George and Mallery (2003) and are broadly consistent with earlier guidelines by Nunnally (1978):

Cronbach’s AlphaInterpretation
≥ 0.90Excellent
0.80 – 0.89Good
0.70 – 0.79Acceptable
0.60 – 0.69Questionable
0.50 – 0.59Poor
< 0.50Unacceptable

The UCLA Statistical Consulting Group’s guide to Cronbach’s alpha provides a thorough explanation of these thresholds with annotated SPSS output and additional worked examples.

Three important cautions with these thresholds:

  1. Too high is also a warning. An alpha above 0.95 often indicates item redundancy — the items are so similar they contribute no unique information. A well-designed scale usually targets 0.80–0.92 rather than maximising alpha.

  2. Alpha inflates with item count. Adding more items raises alpha mechanically, regardless of whether those items improve measurement quality. A 20-item scale can clear 0.80 with only modest inter-item correlations; a 3-item scale needs strong correlations to reach 0.70.

  3. Alpha assumes a unidimensional scale. If a questionnaire secretly measures two distinct sub-constructs, a single alpha computed across all items may be misleadingly moderate even when each sub-dimension is measured well. Run a factor analysis to check dimensionality, then compute alpha separately for each factor if needed.


Inter-Rater Reliability: Cohen’s Kappa

Cohen’s kappa (κ) — sometimes seen as cohen kappa in software output — measures the degree of agreement between two raters who independently assign subjects to discrete categories, correcting for the baseline agreement that would be expected by chance alone.

Consider two clinicians independently classifying chest X-rays as “normal” or “abnormal.” Even if both raters assigned labels randomly, they would agree on roughly half the cases simply by chance. Kappa subtracts that chance-level agreement before assessing whether the observed level of agreement is meaningful. A kappa of 0 means the raters agree exactly as much as chance predicts; a kappa of 1 means perfect agreement; negative values (rare in practice) mean they agree less than chance.

The Cohen’s Kappa Formula

κ = (Pₒ − Pₑ) / (1 − Pₑ)

Where:

  • Pₒ = the observed proportional agreement (the fraction of all cases where both raters assign the same category)
  • Pₑ = the expected proportional agreement by chance, computed from the marginal totals of the rating matrix

For a two-category system with n total cases, Pₑ is:

Pₑ = ((n₁+ × n+₁) + (n₂+ × n+₂)) / n²

Here n₁+ and n₂+ are the row totals for Rater A, and n+₁ and n+₂ are the column totals for Rater B. The formula generalises to any number of categories by summing over all diagonal cells in the same way.

Step-by-Step Worked Example

Two pathologists review 40 tissue samples and independently classify each as “Malignant” or “Benign”:

Path. B: MalignantPath. B: BenignRow Total
Path. A: Malignant20525
Path. A: Benign31215
Column Total231740

Step 1 — Compute observed agreement (Pₒ).

Both say Malignant: 20. Both say Benign: 12. Agreements = 32.

Pₒ = 32 / 40 = 0.800

Step 2 — Compute expected agreement (Pₑ).

Expected joint “Malignant” classification:

  • Rater A’s Malignant rate: 25/40 = 0.625
  • Rater B’s Malignant rate: 23/40 = 0.575
  • Joint expected: 0.625 × 0.575 = 0.35938

Expected joint “Benign” classification:

  • Rater A’s Benign rate: 15/40 = 0.375
  • Rater B’s Benign rate: 17/40 = 0.425
  • Joint expected: 0.375 × 0.425 = 0.15938
Pₑ = 0.35938 + 0.15938 = 0.51875

Step 3 — Apply the kappa formula.

κ = (0.800 − 0.51875) / (1 − 0.51875)
κ = 0.28125 / 0.48125
κ ≈ 0.584

A Cohen’s kappa of 0.58 places the two pathologists in the moderate agreement range — better than chance, but not yet at the level clinical decision-making typically demands.


Interpreting Cohen’s Kappa

The Landis and Koch (1977) benchmarks remain the most widely cited guide for interpreting kappa:

Cohen’s KappaStrength of Agreement
< 0.00Poor (less than chance)
0.00 – 0.20Slight
0.21 – 0.40Fair
0.41 – 0.60Moderate
0.61 – 0.80Substantial
0.81 – 1.00Almost Perfect

The worked example above (κ ≈ 0.58) sits at the upper end of the moderate band, just below substantial.

Context determines whether a given kappa is acceptable. In clinical diagnostics — where a wrong classification can affect treatment — researchers often require κ ≥ 0.80 before using a rating scheme. In early-stage qualitative content analysis for social science research, a kappa in the 0.60–0.70 range is commonly considered adequate while raters calibrate their coding guide.

Prevalence and bias effects. Kappa is sensitive to how evenly cases are distributed across categories. When one category is much rarer than another (high prevalence imbalance), kappa can be misleadingly low even when raw percentage agreement looks high. Researchers sometimes supplement kappa with the prevalence-adjusted, bias-adjusted kappa (PABAK) in those situations. An accessible discussion of these nuances appears in the NIH National Library of Medicine open-access review of interrater reliability and the kappa statistic.


Cronbach’s Alpha vs. Cohen’s Kappa: When to Use Each

The two statistics are not interchangeable. They address entirely different research questions:

ScenarioStatistic to Use
Multiple survey/test items scored on a continuous or ordinal scaleCronbach’s alpha
Two raters independently classifying cases into discrete categoriesCohen’s kappa
You want to report internal consistency for a published psychometric scaleCronbach’s alpha
You need to establish inter-rater agreement before data collection beginsCohen’s kappa
Scale items use Likert responses (1–5, 1–7, strongly agree → strongly disagree)Cronbach’s alpha
Ratings are nominal (yes/no, diagnosis A/B/C, pass/fail)Cohen’s kappa

Applying Cronbach’s alpha to inter-rater data, or Cohen’s kappa to multi-item scales, is a conceptual error — not merely a suboptimal choice. Each statistic measures something the other cannot.

Note also that Cohen’s kappa in its standard form applies to exactly two raters and nominal categories. For ordinal categories (where disagreement by one level is less serious than disagreement by three levels), use weighted kappa. For three or more raters, use intraclass correlation coefficients (ICC) for continuous outcomes or Fleiss’s kappa for nominal ones.


Common Mistakes with Reliability Statistics

Running Cronbach’s Alpha on a Multi-Dimensional Scale

A questionnaire measuring two distinct sub-constructs — for example, physical and emotional symptoms — will produce a single overall alpha that is lower than either sub-scale’s alpha, or misleadingly moderate despite genuinely strong sub-scale reliability. Before computing alpha, run a principal component analysis or confirmatory factor analysis to check whether your items are truly unidimensional. If the scale has multiple factors, compute and report alpha separately for each.

Reporting Kappa Without the Confusion Matrix

A kappa value without the underlying frequency table is hard to evaluate. A kappa of 0.55 in a dataset where 95% of cases fall into one category is very different from the same kappa in a balanced dataset. Always report the full rating matrix alongside kappa so readers can assess the prevalence distribution themselves.

Treating High Alpha as Evidence of Validity

Reliability and validity are separate properties. A questionnaire can achieve Cronbach’s alpha of 0.93 while consistently measuring the wrong construct. High internal consistency confirms that the items cohere — not that they measure what the researcher intends. Validity requires additional evidence: convergent and discriminant validity correlations, factor structure, criterion-related validation studies, and content reviews by domain experts.

Ignoring Sample Size Effects

Both Cronbach’s alpha and Cohen’s kappa are sample estimates and are subject to sampling error. With fewer than 30–50 observations, alpha estimates are unstable and kappa confidence intervals can be extremely wide. Report confidence intervals alongside point estimates wherever your software provides them — they communicate the uncertainty in your reliability estimate that a bare number hides.

Forgetting to Reverse-Score Items Before Computing Alpha

Negatively worded items (e.g., “I rarely feel engaged at work” on an engagement scale) must be reverse-scored before computing alpha. Failing to do this artificially lowers alpha — sometimes to negative values — because the item is pulling in the opposite direction from the rest of the scale. A negative alpha is almost always a sign of an unrecoded reverse-scored item, not a genuinely unreliable scale.


Frequently Asked Questions

What does Cronbach’s alpha measure?

Cronbach’s alpha measures the internal consistency of a multi-item scale — how strongly all items on the scale correlate with each other. High alpha indicates the items are measuring the same underlying construct; low alpha suggests the items tap different constructs or introduce a lot of measurement noise. The statistic ranges from 0 to 1, with 0.70 commonly used as the minimum acceptable threshold in social-science research.

What is an acceptable Cronbach’s alpha?

The most widely cited minimum is 0.70. Values of 0.80 or above indicate good reliability; values of 0.90 or above indicate excellent reliability. However, very high alpha (above 0.95) may signal item redundancy — items that are so similar they add no new measurement information and inflate the scale unnecessarily. The appropriate threshold also depends on the stakes: clinical or diagnostic applications typically demand higher alphas than exploratory pilot research.

What does Cohen’s kappa tell you?

Cohen’s kappa quantifies how much two raters agree beyond the level of agreement that pure chance would produce. A kappa of 0 means the raters agree no more than chance predicts; a kappa of 1 means perfect agreement. Values above 0.60 generally indicate substantial agreement; values above 0.80 indicate near-perfect agreement. Negative kappa values (agreement worse than chance) are theoretically possible but very rare in practice.

What is a good Cohen’s kappa value?

Following the Landis and Koch (1977) benchmarks, kappa values of 0.61–0.80 indicate substantial agreement and are generally considered acceptable for published research. Values of 0.81–1.00 indicate almost perfect agreement. What counts as “good enough” depends heavily on context — a kappa of 0.70 may be fine for content-coding a social-media corpus but unacceptable for classifying pathology slides.

Can Cronbach’s alpha be negative?

Yes. A negative alpha occurs when the sum of item variances exceeds the variance of the total score — which happens when items are negatively correlated with the scale total. In practice this almost always means a reverse-scored item was not recoded before analysis, or that the items are measuring genuinely opposing constructs and should not be combined into a single scale.

How does Cronbach’s alpha relate to a split-half reliability coefficient?

Cronbach’s alpha equals the mean of all possible split-half reliability coefficients of a scale, adjusted for the number of items using the Spearman-Brown prophecy formula. This is why alpha is sometimes described as a generalisation of split-half reliability. Running a single arbitrary split (odd items vs. even items, for example) gives one estimate of the same underlying quantity; alpha averages across all possible splits and is therefore a more stable estimate.

Why does Cohen’s kappa differ from simple percentage agreement?

Simple percentage agreement does not account for the agreement that would occur by chance if both raters assigned categories randomly. In a two-category system where Rater A classifies 80% of cases as “positive,” both raters would agree on about 80% × 80% = 64% of “positive” cases by chance alone. Cohen’s kappa subtracts this expected-by-chance component, so it reflects only the agreement above and beyond what chance produces — a more conservative and informative measure of genuine rater concurrence.


Summary

Reliability is the bedrock of credible measurement. Two statistics carry most of the practical work in quantitative research:

  • Cronbach’s alpha (also called Cronbach alpha) measures internal consistency — how cohesively a set of scale items measure the same underlying construct. Use it when you have multiple items on a survey or test scored on continuous or ordinal scales. An alpha of 0.70 is a common minimum; 0.80 or above is good; 0.90 or above is excellent (though values above 0.95 may indicate item redundancy).

  • Cohen’s kappa measures inter-rater reliability — how much two independent raters agree on categorical classifications beyond what chance would produce. Use it when two observers independently assign discrete labels to the same set of cases. Values above 0.60 indicate substantial agreement; values above 0.80 indicate near-perfect agreement. Interpret kappa alongside its confidence interval and the full rating matrix.

The two statistics are not interchangeable and should never be applied to each other’s domain. Cronbach’s alpha operates on item-level scores within a single instrument; Cohen’s kappa operates on labels assigned by separate observers. Keeping these distinct in both your analysis plan and your written report prevents a common and consequential methodological error.

Both statistics have known limitations — alpha assumes unidimensionality; kappa is sensitive to category prevalence — and both benefit from being reported with confidence intervals and contextual explanation rather than as bare numbers against a universal threshold.