Mediation analysis measures how much of the relationship between an independent variable (X) and an outcome (Y) flows through a third variable, the mediator (M), rather than acting directly. The practical recommendation from decades of methods research is straightforward: estimate the indirect effect as the product of two path coefficients, then report a bootstrapped confidence interval instead of relying on older significance tests. If you want to make a causal claim rather than a descriptive one, you also need to check a specific set of confounding assumptions before you trust the number.
Key takeaways
| Point | Details |
|---|---|
| Define your paths first | Label a (X to M), b (M to Y), c (total effect), and c' (direct effect) before running any test. |
| Skip the Sobel test | Bootstrap the indirect effect instead, since the product ab has a skewed sampling distribution. |
| Significance doesn't require a big total effect | An indirect effect can be real even when c is close to zero, since paths can offset each other. |
| Causal claims need DAG-checked assumptions | Confirm no unmeasured X-M, M-Y, or collider confounding, and run a sensitivity analysis. |
| Build the surrounding skills first | Statohub's Experiments & Causality hub and calculators support the causal reasoning and descriptive work mediation analysis depends on. |
This page is built for readers who need to run this analysis, not just recognize the term. By the end, you will be able to:
- Write the regression equations behind a single-mediator model and label each path correctly
- Compute an indirect effect and interpret a bootstrapped confidence interval
- Recognize when your design supports a causal claim and when it does not
What Is Mediation Analysis and How Does It Differ From Moderation?
Mediation analysis is a statistical method for identifying and quantifying the mechanism connecting X to Y through an intervening variable, as MacKinnon and colleagues define it in their widely cited review. Picture a simple chain: a training program (X) increases an employee’s confidence (M), and that added confidence increases performance ratings (Y). The mediator is not just a bystander variable. It is the mechanism, the “how,” that explains why X and Y are connected in the first place.
The single-mediator model rests on three regression equations, each producing a labeled path:
- Y regressed on X alone: Y = i₁ + cX + e₁, where c is the total effect of X on Y.
- M regressed on X: M = i₂ + aX + e₂, where a is the effect of X on the mediator.
- Y regressed on both X and M: Y = i₃ + c’X + bM + e₃, where b is the effect of the mediator on Y controlling for X, and c’ is the direct effect of X on Y once M is accounted for.
From these three equations, you get everything you need. The indirect effect is a times b, the portion of X’s influence on Y that travels through the mediator. The direct effect is c’, what remains of X’s influence on Y after removing the mediated pathway. The total effect, c, is the sum of the two: c = c’ + ab. When c’ shrinks to near zero once M enters the model, researchers call this full mediation. When c’ stays meaningfully nonzero, that is partial mediation, meaning the mediator explains part of the story but not all of it.
Confusing a mediator with a moderator is one of the most frequent errors in applied research, and it changes the entire analysis. A mediator explains the mechanism between X and Y. A moderator changes the strength or direction of that relationship depending on its own value, acting as a boundary condition rather than a pathway. The University of Southampton’s library guide puts it plainly: ask “how does X affect Y?” and you are looking for a mediator; ask “for whom, or under what conditions, does X affect Y?” and you are looking for a moderator. A confounder is different still: it is a common cause of both X and Y that was never part of the causal chain you are trying to describe, and failing to distinguish it from a mediator can quietly invalidate your conclusions.
How Do You Compute an Indirect Effect Mathematically?
Two mathematically distinct routes lead to the same number under ordinary linear regression, and understanding both builds real intuition for what mediation analysis is doing under the hood.
The product-of-coefficients method multiplies the X→M path (a) by the M→Y path (b), giving you ab directly from your two regression models. The difference-in-coefficients method instead compares the total effect to the direct effect: c minus c’. Under standard ordinary least squares assumptions with a continuous outcome, these two quantities are algebraically identical: ab = c − c’. That equivalence is not a coincidence. It falls directly out of substituting the M-on-X equation into the Y-on-X-and-M equation and simplifying terms.
The equivalence breaks down once you leave simple linear regression. In multilevel models or with binary outcomes fit through logistic regression, ab and c − c’ can diverge because the transformations involved are not linear, a point the PMC review by MacKinnon, Fairchild, and Fritz discusses in detail.
For decades, the standard significance test for the indirect effect was the Sobel test, which divides ab by its standard error to produce a z-score. The formula works, but it rests on an assumption that rarely holds: that the sampling distribution of ab is normal. In reality, the product of two normally distributed coefficients is skewed, often heavily, especially with smaller samples. That mismatch between assumption and reality is the main reason the Sobel test tends to underpower studies and produce confidence intervals that don’t actually cover the true value at the stated rate, a shortfall documented directly in a head-to-head comparison of mediation test methods (MacKinnon, Lockwood, Hoffman, West & Sheets, 2002).
A finding that reshaped applied practice: an indirect effect can be statistically significant even when the total effect (c) is nowhere near significance, because direct and indirect effects sometimes point in opposite directions and partially cancel each other out. This single insight, documented in the PMC mediation review, is why the older Baron and Kenny requirement of a significant total effect before testing mediation is no longer treated as a prerequisite.
Beyond significance, report an effect-size measure so readers can judge practical importance, not just statistical detectability. Common choices include:
- The proportion mediated (ab divided by c), interpreted cautiously since it can behave oddly near zero
- A standardized indirect effect, which rescales a and b into comparable units
- Completely standardized coefficients when your X, M, and Y are measured on different scales
Why Should You Bootstrap the Indirect Effect Instead of Using the Sobel Test?
Because the sampling distribution of ab is asymmetric, not bell-shaped, and normal-theory tests like Sobel’s assume symmetry that simply is not there. Bootstrapping sidesteps the assumption entirely by building the distribution of ab empirically from your own data, which is why it has become the default recommendation in modern mediation practice rather than a niche alternative.
The logic behind the bootstrap recipe below is simple enough to run by hand conceptually, even though you’ll do it in software, and it follows the same resampling structure described in Preacher and Hayes’s widely used SPSS and SAS procedures for estimating indirect effects.
Two flavors of the resulting interval show up in the literature and in software output: the percentile method, which takes the raw percentiles directly from the bootstrap distribution, and the bias-corrected and accelerated (BCa) method, which adjusts for skew and bias in that distribution. BCa intervals are generally preferred when the bootstrap distribution is noticeably asymmetric, which is common with smaller samples or weak paths.
A few practical notes worth keeping in mind before you run this yourself:
- Mediation effects typically need larger samples than simple two-group comparisons because you are estimating the precision of a product of two coefficients, not just one.
- Rough guidance from simulation studies suggests that detecting small indirect effects reliably often calls for samples in the hundreds rather than dozens, particularly when both a and b paths are individually modest.
- Power drops sharply when either the a path or the b path is weak, even if the other path is strong. A weak link anywhere in the chain limits the whole product.
For implementation, R users generally reach for the mediation package or structural equation modeling through lavaan, both of which build bootstrapped confidence intervals natively. SPSS users typically run Andrew Hayes’s PROCESS macro, which the IBM SPSS documentation also references directly, and which produces bootstrapped indirect-effect estimates without requiring the user to code the resampling loop manually. Python users can replicate the same logic using statsmodels for the underlying regressions paired with a manual resampling loop, or dedicated packages built for causal mediation.
What Assumptions Does Mediation Analysis Need to Support Causal Claims?
Running the regressions correctly is the easy part. The hard part, and the part most textbooks underplay, is the set of untestable assumptions standing between your coefficients and a genuine causal claim. As one methodological perspective on confounding puts it, the real difficulty in mediation is not computational but assumption-driven: the math works regardless of whether your causal story is right.
Four assumptions need to hold, and none of them can be verified from the data alone:
- No unmeasured confounders affecting both X and M (the X→M relationship must be free of common causes you haven’t measured).
- No unmeasured confounders affecting both M and Y (the M→Y relationship needs the same clean status).
- No unmeasured confounders affecting both X and Y outside the mediated pathway (this overlaps with standard confounding concerns in any causal claim about X and Y).
- No mediator-outcome confounder that is itself caused by X (this is the subtlest one, and it trips up even careful analysts).
That last condition deserves a moment, because it is where good intentions produce bad statistics. Suppose X is a diet intervention, M is inflammation, and Y is disease risk. If you control for a variable, say medication use, that is itself affected by X and also affects both M and Y, you have conditioned on a collider. Doing so can introduce a spurious association between M and Y that has nothing to do with the true mechanism, distorting your b path in ways that are hard to detect just by looking at p-values.
Directed acyclic graphs, or DAGs, are the standard tool for reasoning through this before you touch the regression. Drawing out every variable you believe influences X, M, or Y, and the arrows connecting them, forces you to be explicit about which variables belong in your model as controls and which ones will corrupt it if included. Statohub’s guide to correlation versus causation covers the broader logic of why an association alone never proves a causal story, a caution that applies with particular force here.
Because these confounding assumptions cannot be tested directly, sensitivity analysis is the practical substitute. The general approach asks: how strong would an unmeasured confounder need to be, in terms of its association with both the mediator and the outcome, before it would explain away your observed indirect effect? If the answer is “implausibly strong,” your finding is more robust. If a modest, plausible confounder could erase the effect, you should say so plainly in your write-up.
How Does Mediation Get More Complicated With Multiple Mediators?
Real processes rarely funnel through exactly one mechanism, and mediation models scale up to reflect that in a few standard ways.
Parallel mediators sit side by side, each with its own independent a-path from X and b-path to Y, with no arrows connecting the mediators to each other. The total indirect effect becomes the sum of each mediator’s individual ab product, letting you see how much of the total mechanism each pathway accounts for. Serial mediators, by contrast, form a chain, where X affects M1, M1 affects M2, and M2 affects Y. The indirect effect through a serial chain multiplies across every link in the sequence, and specific indirect effects for each possible path through the chain can be extracted and reported separately.
Moderated mediation asks whether the strength of an indirect effect itself depends on a third variable. Rather than a single ab value, you get a conditional indirect effect that changes across levels of the moderator. The index of moderated mediation formalizes this: it tests whether the a path, the b path, or both, vary systematically with the moderator, giving you a single coefficient and confidence interval that answers “does the mechanism strengthen or weaken depending on this other variable?” rather than forcing you to eyeball separate models at different moderator values.
Mediation with binary outcomes or multilevel, nested data introduces real complications that the simple linear framework glosses over:
- Logistic regression coefficients for a binary Y are on the log-odds scale, and multiplying an a path from a linear model with a b path from a logistic model does not produce an interpretable quantity without additional transformation.
- Counterfactual and causal mediation frameworks, like those detailed in Columbia University’s causal mediation resources, handle these non-linear cases more rigorously by defining natural direct and indirect effects through potential outcomes rather than raw coefficients.
- In multilevel data, where individuals are nested within groups such as classrooms or clinics, the difference-in-coefficients and product-of-coefficients methods are no longer guaranteed to match, requiring either multilevel structural equation modeling or Monte Carlo methods for the indirect effect’s confidence interval.
- If X interacts with M in predicting Y, the simple ab product can misrepresent the average indirect effect, and analysts need interaction-aware formulas or counterfactual estimands, a point raised in UCLA’s causality research notes.
A Step-by-Step Worked Example You Can Replicate
Consider a study of 250 employees examining whether a time-management training program (X) with control and treatment groups reduces reported burnout (Y, measured on a validated 20-point scale) through improved perceived control over one’s workload (M, measured on a validated 15-point scale). This mirrors the structure recommended in the University of Virginia Library’s mediation walkthrough, which frames mediation instruction around exactly this kind of three-regression sequence.
Fitting the three regressions described earlier on this hypothetical sample produces the following illustrative coefficients:
| Path | Coefficient | Description |
|---|---|---|
| a (X → M) | 2.10 | Training raises perceived control by 2.10 points on average |
| b (M → Y, controlling for X) | −0.45 | Each point of added control is associated with a 0.45-point drop in burnout |
| c (X → Y, total effect) | −1.35 | Training alone is associated with a 1.35-point drop in burnout |
| c' (X → Y, direct effect) | −0.40 | Direct effect shrinks substantially once M is added |
| ab (indirect effect) | −0.945 | 2.10 × (−0.45) |
Here, the total effect (−1.35) splits into a direct effect (−0.40) and an indirect effect (−0.945), and −0.40 plus −0.945 lands close to −1.35, confirming the algebra holds within rounding.
The next step is inference on that −0.945 indirect effect. Because that interval excludes zero entirely, the indirect effect through perceived control is distinguishable from noise at conventional confidence levels, even though the direct effect of −0.40 might not clear significance on its own in a smaller sample.
Reporting sentences for a result like this typically follow a pattern:
- “Time-management training was associated with a total reduction in burnout of 1.35 points (c = −1.35).”
- “This effect operated substantially through increased perceived control over workload (a = 2.10, b = −0.45, indirect effect ab = −0.945, 95% BCa CI [−1.42, −0.51]).”
- “The direct effect of training on burnout after accounting for perceived control was smaller (c’ = −0.40), consistent with partial mediation.”
Statohub’s effect size guide is a useful companion here for translating that −0.945 coefficient into a standardized metric your readers outside the immediate field can interpret.
What Belongs in a Mediation Analysis Report?
A complete mediation write-up gives readers everything they’d need to reproduce your conclusion, not just your headline finding. At minimum, include:
- All three unstandardized path coefficients (a, b, c, c’) with their standard errors
- The indirect effect (ab) and its bootstrapped confidence interval, specifying percentile or BCa
- Sample size, number of bootstrap resamples, and whether you set a random seed
- A path diagram showing X, M, and Y with coefficients labeled on each arrow
When you write up the interpretation, resist collapsing the finding into “M mediates the relationship between X and Y” without qualification. State the direction, the magnitude, and at least one plausible alternative explanation the design cannot rule out, whether that’s reverse causation, an unmeasured confounder, or measurement error in the mediator. A mediation result is a piece of evidence, not a verdict, and framing it that way protects your credibility more than it costs you rhetorical punch.
For reproducibility, share your analysis code, the random seed used for bootstrapping, and the exact number of resamples run. This matters more here than in ordinary regression, because bootstrap results shift slightly from run to run without a fixed seed, and reviewers or replicators need to know whether small numeric differences reflect a real disagreement or just resampling variation.
What Mistakes Undermine a Mediation Analysis Most Often?
The most common failure mode is treating the Baron and Kenny causal-steps approach as sufficient on its own, without ever computing a confidence interval for the indirect effect itself. Causal steps can tell you a pattern of significance exists across three regressions, but they never quantify the indirect effect or its uncertainty, which is the number readers actually need.
Watch for these recurring issues:
- Reverse causation in the mediator. If Y could plausibly cause M rather than the reverse, your b-path is measuring the wrong direction entirely, and no amount of bootstrapping fixes a mis-specified causal order.
- Measurement error in the mediator. An unreliable mediator scale attenuates the b path toward zero, which can make a real mediation effect look weaker than it actually is or disappear altogether.
- Conditioning on a collider. Adding a control variable that is itself caused by X and also affects M or Y introduces bias that looks like a legitimate covariate adjustment but isn’t.
- Cross-sectional data measuring X, M, and Y at the same timepoint. Without temporal separation, you cannot rule out that Y influenced M rather than the other way around.
When these diagnostics raise red flags, the honest move is not to abandon the analysis but to adjust your claims to match your evidence. Run the sensitivity analysis described earlier and report how robust the effect is to a plausible confounder. Soften causal language to something like “consistent with” rather than “demonstrates.” And where the research question genuinely warrants it, consider whether an experimental manipulation of the mediator itself, rather than a purely observational design, is feasible in a follow-up study, since directly intervening on M is one of the few ways to test a mediating mechanism with real causal leverage.
Where Statohub Fits Into Learning This Method
Mediation analysis sits at the intersection of regression, causal reasoning, and inference, which is exactly the territory Statohub’s Experiments & Causality hub is built to cover. If the DAG reasoning in the assumptions section above felt new, that hub walks through confounding, randomization, and causal interpretation in more depth, using the same practical framing applied here.
For the computational side, Statohub’s calculator library supports the descriptive and inferential work that typically comes before or alongside a mediation model, from checking distributional assumptions to computing basic averages your regressions depend on. Pair that with the effect size guide once you’ve computed your ab coefficient and want to express it in standardized terms a broader readership can interpret.
The learn, calculate, apply structure exists precisely for readers moving from “I understand the concept” to “I can produce this analysis myself.” Mediation analysis is a strong test case for that flow, since it demands conceptual clarity about paths and assumptions, computational comfort with regression and bootstrapping, and applied judgment about when a causal claim is actually warranted.
Choosing Instruments for Your Mediator and Outcome
An indirect effect is only as trustworthy as the instruments used to measure M and Y, and this is where many otherwise well-designed mediation studies quietly fail. A mediator scale with weak internal consistency, a Cronbach’s alpha well below 0.70, for instance, introduces measurement error that attenuates the b path toward zero and can hide a real mediating mechanism entirely.
Before committing to an instrument, check for two properties separately. Reliability asks whether the measure produces consistent scores across repeated administrations or across items intended to tap the same construct. Validity asks whether the instrument actually measures the construct you think it does, rather than something adjacent to it. A workload-control scale that actually captures general job satisfaction, for example, will produce a mediation result that looks clean statistically but answers the wrong substantive question.
Where possible, favor validated instruments with published psychometric properties over ad hoc scales built for a single study. This matters even more for the mediator than the outcome, since a noisy M distorts both the a path and the b path simultaneously, compounding the error in your indirect effect estimate. When you cannot avoid a newly developed measure, report its reliability statistics alongside your mediation results so readers can judge how much attenuation might be affecting your conclusions.
A Practical Checklist for Running Your Own Mediation Study
Design decisions made before you collect a single data point matter more than any modeling choice made afterward, and this is where Statohub’s editorial view diverges from how many methods courses sequence the material. Most curricula teach the Sobel test, then the bootstrap, then assumptions, as though causal reasoning were an afterthought bolted onto the statistics. Reverse that order in practice.
Practical checklist for running your own mediation study
- Draw your DAG before collecting data Specify every variable you believe influences X, M, or Y before you touch the regression.
- Choose validated, reliable instruments for M and Y Pilot them first if they're newly developed.
- Establish temporal separation Measure X, M, and Y at different timepoints whenever your design allows it.
- Fit the three regressions Compute a, b, c, and c' from the X-on-M, Y-on-X, and Y-on-X-and-M models.
- Bootstrap the indirect effect Use at least several thousand resamples and extract a BCa interval.
- Run a sensitivity analysis Check how robust the indirect effect is to unmeasured confounding on the M-to-Y path.
- Write up direct, indirect, and total effects Report confidence intervals for each one, not just significance stars.
If your bootstrap interval is wide enough to include values near zero and values several times larger, that is a signal to collect more data before publishing a confident interpretation, not a cue to round the finding up to “significant” and move on. For classroom practice, a strong exercise is replicating a published mediation finding from a paper in your field using only the reported coefficients, then checking whether your recomputed indirect effect and interval match what the authors reported.
Put These Mediation Methods Into Practice
The regression math behind mediation is approachable once you’ve worked through the three equations above, but running the bootstrap and checking your assumptions by hand each time is where most people stall out. Statohub’s Applied Statistics hub exists to close that gap, with guides that walk through causal reasoning and experimental design the same way this page walked through path coefficients and confidence intervals.
If you want to check your descriptive statistics before you build out a mediation model, Statohub’s calculators handle the groundwork, Statohub’s Data Analysis guides cover that descriptive workflow in more depth, and the Experiments & Causality section goes deeper into the confounding and DAG logic that determines whether your indirect effect deserves a causal interpretation at all. Start there, work through a real dataset from your own coursework or research, and compute your first bootstrapped indirect effect using the steps outlined above.
Sources
Sources
- MacKinnon DP, Fairchild AJ, Fritz MS — "Mediation Analysis," Annual Review of Psychology (2007) PubMed Central / NIH
- MacKinnon DP, Lockwood CM, Hoffman JM, West SG, Sheets V — "A Comparison of Methods to Test Mediation and Other Intervening Variable Effects," Psychological Methods (2002) PubMed Central / NIH
- Preacher KJ, Hayes AF — "SPSS and SAS Procedures for Estimating Indirect Effects in Simple Mediation Models," Behavior Research Methods (2004) PubMed / National Library of Medicine
- Baron RM, Kenny DA — "The Moderator-Mediator Variable Distinction in Social Psychological Research," Journal of Personality and Social Psychology (1986) PubMed / National Library of Medicine
- Introduction to Mediation Analysis | University of Virginia Library University of Virginia
- Mediation vs. Moderation | University of Southampton Library University of Southampton
- Causal Mediation | Population Health Methods Columbia University Mailman School of Public Health
- Comments on Kenny's Summary of Causal Mediation UCLA Causality Lab
FAQ
Frequently asked questions
- Is mediation analysis the same as testing for a confounder?
- No. A confounder is a common cause of X and Y that sits outside the causal chain you're studying, while a mediator is part of that chain, explaining the mechanism through which X influences Y. Treating a confounder as a mediator, or vice versa, leads to fundamentally different modeling choices and conclusions.
- Can I run mediation analysis on cross-sectional data?
- You can compute the coefficients, but a causal mediation interpretation is weaker without temporal separation between X, M, and Y. Cross-sectional designs can't rule out reverse causation between the mediator and outcome, so treat cross-sectional mediation results as suggestive rather than confirmatory.
- How many bootstrap resamples should I use?
- Several thousand resamples is a common default that balances stability and computation time for most academic and applied work. Some analysts use 10,000 for published results, though the confidence interval typically stabilizes well before that point.
- What sample size do I need for mediation analysis?
- It depends heavily on the strength of your a and b paths individually, not just their product. Detecting small effects reliably in the ab product generally requires moderately large samples to achieve sufficient power, since you're estimating uncertainty around a product of two coefficients.
- Does a nonsignificant total effect mean there's no mediation?
- No, and this is one of the more counterintuitive findings in the field. Direct and indirect effects can run in opposite directions and cancel out, producing a near-zero total effect while a genuine, statistically distinguishable indirect effect exists underneath it.