A normality test checks whether a sample plausibly came from a normal distribution, using a formal statistic and p-value rather than a visual guess. For most student and research work, start with the Shapiro–Wilk test, add Anderson–Darling when tail behavior matters, and use D’Agostino or Jarque–Bera for a quick skewness and kurtosis screen. Whatever you pick, remember the central trap: failing to reject the null hypothesis never proves your data are normal.
Key takeaways
| Point | Details |
|---|---|
| Match the test to sample size | Shapiro–Wilk is most powerful for small samples under 50, while Anderson–Darling better detects tail deviations; large samples call for effect-size checks over raw p-values. |
| Q–Q plots are the anchor visual | They are the most reliable visual check; curvature at the ends signals heavy tails or skewness, especially in larger datasets. |
| Large-sample p-values can mislead | Shapiro–Wilk and Anderson–Darling both grow more sensitive at scale, so a small p-value on n > 1,000 often reflects a trivial, practically meaningless difference. |
| Failing to reject is not proof of normality | A p-value above 0.05 only means you lack evidence against normality, not that the data are normal — treat it as evidence, not proof, and confirm with more than one check. |
| Have a fallback plan | When data are genuinely non-normal, a transformation or a nonparametric test is usually a better fix than forcing a parametric assumption. |
Quick Checklist: Which Normality Test to Run
Before touching a dataset, match the test to the sample and the question you actually need answered. The right choice depends less on habit and more on sample size, tail sensitivity, and how the parameters were estimated.
Quick checklist: which normality test to run
- Small samples (roughly n < 50) Favor Shapiro–Wilk — it consistently shows strong power in this range, per comparative reviews.
- Tail or extreme-value concerns Run Anderson–Darling, which weights the tails of the distribution more heavily than other empirical distribution function (EDF) tests.
- Quick omnibus screen D'Agostino's test or Jarque–Bera flag unusual skewness or kurtosis fast, though they miss some subtler departures from normality.
- CDF-distance test with estimated parameters Use the Lilliefors-adjusted Kolmogorov–Smirnov test rather than the classic KS test, which assumes known parameters.
- Always Pair whichever statistical test for normality you choose with a Q–Q plot and a look for outliers before trusting the number alone.
Think of this as triage, not dogma. A single test rarely settles the question on its own.
What Do Histograms, Density Plots, and Q–Q Plots Show You?
A Q–Q plot is usually the most reliable visual check, and it is worth learning to read well. It plots your sample’s ordered values against the quantiles a normal distribution would produce, and points that fall close to a straight diagonal line indicate a roughly normal shape. Curvature at the ends signals heavy or light tails, while an S-shape suggests skewness. NIST’s exploratory data analysis handbook recommends probability plots over histograms precisely because histograms depend on arbitrary bin widths and need a fairly large sample before their shape becomes trustworthy.
Density plots smooth over some of that binning noise, but they can still mislead with small samples by exaggerating minor bumps as if they were real modes. Boxplots earn their place here too: they flag outliers quickly through points beyond the whiskers, and an off-center median relative to the box hints at skew before you run a single test.
Watch for a few recurring traps. Rounded or heavily tied data (survey responses on a 1 to 5 scale, for instance) can make a Q–Q plot look stair-stepped rather than smooth, which is a data artifact, not necessarily evidence against normality. Small samples add their own visual noise, so a slightly wavy Q–Q plot with n = 15 deserves less alarm than the same wave with n = 500.
What Do Shapiro–Wilk, Anderson–Darling, and Other Normality Tests Measure?
Each core normality test approaches the same question from a different angle, and knowing the mechanics helps you pick correctly rather than by habit.
| Test | What it measures |
|---|---|
| Shapiro–Wilk | Q–Q straightness; a W statistic close to 1 indicates a good fit to normality. |
| Anderson–Darling | Empirical CDF vs. the normal curve, with extra weight applied to the tails. |
| Kolmogorov–Smirnov (Lilliefors adjusted) | Maximum vertical distance between the sample and theoretical normal CDF, adjusted for estimated parameters. |
| D'Agostino / Jarque–Bera | Sample skewness and kurtosis against what a normal distribution would produce; a fast omnibus screen. |
| Energy / ECF tests | Joint (multivariate) normality across several variables at once, rather than one variable at a time. |
Shapiro–Wilk computes a statistic, W, from a weighted combination of the ordered sample values. Shapiro and Wilk’s original 1965 paper frames W as essentially a numeric measure of how straight the sample’s Q–Q plot is. Values of W close to 1 indicate a good fit to normality, and the test tends to hold strong power across many small-to-moderate sample settings.
Anderson–Darling belongs to the empirical distribution function family, comparing the sample’s cumulative distribution to the normal curve, but it applies extra weight to the tails. That makes it the better pick when you specifically care about extreme values, such as flagging heavy-tailed risk data.
Kolmogorov–Smirnov, with the Lilliefors adjustment for estimated parameters, measures the maximum vertical distance between the sample’s empirical CDF and the theoretical normal CDF. It has a long history and remains useful for teaching, but it generally has less power than Shapiro–Wilk or Anderson–Darling for detecting tail-driven departures, according to the same comparative literature.
D’Agostino’s test and Jarque–Bera work differently again, checking sample skewness and kurtosis against what a normal distribution would produce. They run fast and give an intuitive omnibus read, but both can miss departures that do not show up cleanly in the third and fourth moments.
Beyond these staples, modern alternatives like energy tests and empirical characteristic function (ECF) based tests extend normality checking into multivariate settings, where you are testing whether several variables jointly follow a multivariate normal distribution rather than checking one variable at a time.
How Should Sample Size and Test Power Shape Your Choice?
Sample size changes not just which test you can run, but how much you should trust the result. The following rules of thumb reflect how test power behaves across common sample ranges.
| Sample size | Recommended approach | Why |
|---|---|---|
| n < 50 | Shapiro–Wilk | Its power advantage is most pronounced here, and it remains valid down to very small n. |
| 50–500 | Shapiro–Wilk or Anderson–Darling | Both perform well; choose Anderson–Darling if tail accuracy matters more than overall fit. |
| n > 1,000 | Q–Q plot and effect size over the p-value alone | Test sensitivity climbs sharply, so even trivial, practically meaningless deviations can produce a statistically significant result. |
| Ties or heavily rounded data | Simulation-based checks | Heavy rounding erodes test power because it artificially compresses the distribution's apparent variability. |
| Any borderline result | Run more than one test | Consistent conclusions across Shapiro–Wilk, Anderson–Darling, and a Q–Q plot carry far more weight than any single number. |
What Does a Normality Test P-Value Actually Tell You?
A normality test’s null hypothesis states that the data come from a normal distribution. The p-value tells you how likely you would be to see a discrepancy this large (or larger) if that null hypothesis were true. A small p-value, typically below 0.05, counts as evidence against normality; it does not measure how normal your data are, only how compatible they look with a perfectly normal population.
The reverse interpretation trips up far more people. Getting p > 0.05 means you failed to reject the null, which is not the same as proving the data are normal. Failing to reject normality often reflects low statistical power rather than genuine normality, especially in small samples where the test simply lacks the sensitivity to detect real departures, according to Statistics By Jim’s guidance on interpreting normality tests.
Ghasemi & Zahediasl, International Journal of Endocrinology and MetabolismFor small sample sizes, normality tests have little power to reject the null hypothesis and therefore small samples most often pass normality tests. For large sample sizes, significant results would be derived even in the case of a small deviation from normality, although this small deviation will not affect the results of a parametric test.
Sample size cuts both ways here. With a very large n, even a skewness of 0.1 or a kurtosis barely off 3 can generate a tiny p-value, flagging “non-normality” that has no practical consequence for your analysis. Always check the effect size, not just the significance flag.
When reporting results, include:
- The test name and statistic (for example, “Shapiro–Wilk, W = 0.97”)
- The p-value and sample size
- A Q–Q plot alongside the numbers, so readers can judge shape directly
What to Do When Your Data Fail a Normality Test
Finding non-normal data is not a dead end. Several honest paths forward exist, and the right one depends on why the data deviate and what your analysis needs.
- Transform the variable. Log, square-root, or Box–Cox transformations often work well for positive, right-skewed data like income or reaction times.
- Switch to robust methods. Bootstrap confidence intervals, trimmed means, and robust regression reduce the influence of outliers and heavy tails without discarding data.
- Move to nonparametric tests. The Mann–Whitney U test, Kruskal–Wallis test, or permutation tests replace their parametric counterparts when the normality assumption breaks down badly.
- Check the model, not just the data. If regression residuals fail a normality check, the problem might be an omitted variable or wrong functional form rather than something a transformation can fix; see this guide to regression assumptions for how to diagnose that distinction.
Skip transformation for its own sake. If your sample size is large and your downstream test is robust to moderate non-normality, the practical cost of non-normality may be smaller than it looks on paper.
A Worked Example: From Histogram to Decision
Picture a dataset of 42 customer wait times, in minutes, collected to check whether an average-based service level agreement is appropriate. The goal is simple: decide whether a parametric method (a t-test on the mean) is defensible, or whether a nonparametric approach makes more sense.
- Look first. A histogram shows a long right tail, with most waits clustered under 10 minutes but a handful stretching past 30. The Q–Q plot confirms it: points curve sharply upward at the top end, consistent with right skew.
- Run two tests. Shapiro–Wilk returns W = 0.89, p = 0.002. Anderson–Darling agrees, rejecting normality with a similarly small p-value. Both point the same direction.
- Inspect outliers. Three wait times above 25 minutes are driving much of the tail. They are real customer records, not data entry errors, so they stay in the dataset.
- Decide and document. With consistent test results, a strongly skewed Q–Q plot, and real (not erroneous) outliers, the honest choice is either a log transformation before running a t-test, or a Mann–Whitney U test if comparing two groups directly.
The lesson generalizes past wait times: consistent evidence across visuals and tests should drive a clear decision, while mixed evidence calls for transparency about which method you chose and why.
Running Normality Tests in R, Python, and SPSS
Each major tool implements these tests with its own quirks worth knowing before you trust the output. In R, shapiro.test() handles sample sizes between 3 and 5,000; beyond that range, it relies on Royston’s approximation, and results should be read with a bit more caution. In Python, SciPy’s stats.shapiro() returns the W statistic and p-value directly, but the SciPy documentation itself warns that p-value accuracy degrades for very large samples, which matters given how sensitive tests become at scale. SPSS reports normality diagnostics along with skewness and kurtosis tables in its Explore procedure, and reading those alongside the test statistic gives a fuller picture than the p-value alone. When your sample size sits outside a tool’s tested range, simulating the null distribution yourself is a safer bet than trusting default output blindly.
Statohub’s Take on Checking for Normality
Statohub’s position is simple: never let a single p-value make the call. Run the test, plot the Q–Q, inspect outliers, and write down what you saw at each step. That habit protects you from both false alarms and false comfort, in a classroom project or a published analysis.
Put These Checks to Work With Statohub’s Tools
Reading about the Shapiro–Wilk test is one thing; watching how a skewed dataset behaves under different transformations is another. Statohub’s Applied Statistics hub walks through hands-on examples that connect these normality checks to real analysis decisions, from regression diagnostics to survey data. If you want a refresher on the distribution these tests measure against, the normal distribution explainer breaks down the bell curve in plain terms, and the Data Analysis hub covers the broader exploratory workflow these tests fit into. For practicing the underlying calculations by hand before you automate them, LogicExcel’s standard deviation exercises offer a useful drill. When you are ready to run the numbers, browse Statohub’s calculators and start checking your own dataset today.
Sources
Sources
- Ghasemi A, Zahediasl S — "Normality Tests for Statistical Analysis: A Guide for Non-Statisticians," International Journal of Endocrinology and Metabolism (2012) PMC / NCBI
- Shapiro, S. S. and Wilk, M. B. — "An Analysis of Variance Test for Normality (Complete Samples)," Biometrika, Vol. 52 (1965) Biometrika
- Shapiro-Wilk test for normality — SciPy Manual SciPy
- NIST/SEMATECH e-Handbook — Normal Probability Plot NIST
- NIST/SEMATECH e-Handbook — Anderson-Darling Test NIST
- NIST/SEMATECH e-Handbook — Anderson-Darling and Shapiro-Wilk Tests NIST
FAQ
Frequently asked questions
- When Should I Use Shapiro–Wilk Instead of Kolmogorov–Smirnov?
- Use Shapiro–Wilk for small-to-moderate samples where you want the most statistical power, since it consistently outperforms Kolmogorov–Smirnov in comparative studies. Reserve the Lilliefors-adjusted Kolmogorov–Smirnov test for cases where you specifically need a CDF-distance measure with estimated parameters, though expect it to detect tail differences less reliably.
- How Do I Interpret a Normality Test Result?
- Read the p-value as evidence against the null hypothesis of normality, not as a normality score. A small p-value (below 0.05) suggests the data likely deviate from normal, while a larger p-value simply means you lack strong evidence of a deviation, which is not the same as proof of normality.
- What Is the Shapiro–Wilk Test for Normality?
- The Shapiro–Wilk test calculates a statistic, W, from the ordered values in your sample, essentially measuring how closely those values align with what a Q–Q plot against a normal distribution would show. Values of W near 1 support normality, and the test tends to hold strong power for small and moderate sample sizes.
- What Does a P-Value of 0.05 Mean in a Shapiro–Wilk Test?
- A p-value at or below 0.05 in a Shapiro–Wilk test signals that the observed data would be unlikely under a true normal distribution, so you reject the assumption of normality at conventional significance levels. That threshold is a convention, not a hard rule, so pair it with a Q–Q plot before making a final call, particularly with large samples where trivial deviations can still produce a low p-value.
- Can I Trust a Normality Test on a Very Large Sample?
- Treat the p-value with caution once your sample climbs past roughly 1,000 observations, because even negligible, practically irrelevant deviations from normal can register as statistically significant. Check skewness, kurtosis, and the Q–Q plot for practical significance rather than relying on the test statistic alone.