Statohub Browse calculators
Data Analysis Practitioner guide

30 Observation Rule: Histogram vs Boxplot, When to Combine

Histogram vs boxplot: use the 30-observation rule to pick the right chart, know which settings to record, and see when combining both works best.

By Statohub Editorial Team Published September 2026Reviewed September 202612 min read

Use a histogram when you need to see distribution shape, modality, and skewness in a single variable. Use a boxplot when you need compact comparisons across groups, along with medians, interquartile ranges, and flagged outliers. Combine both when a report needs shape detail and a concise summary at once, which is common enough in coursework and client decks that it deserves its own workflow below.

Key takeaways

Point Details
Histograms show shape They reveal modality, skewness, and gaps in one variable — but the picture changes with bin width, so rerun the chart at two or three settings before trusting a conclusion about shape.
Boxplots compress to five numbers Median, Q1, Q3, and 1.5×IQR whiskers make many groups comparable at a glance, but a boxplot cannot reveal a second peak in the data.
Below roughly 30 observations, both charts get shaky Bin heights and quartile positions swing wildly with a couple of extra points — plot the raw points instead.
Combine them when you need both A histogram with a marginal boxplot strip gives shape and the five-number summary in one figure without doubling the space.
Always report your settings State the bin rule, the sample size per group, and the outlier rule you used so a reader — or a grader — can reproduce the chart.

What a Histogram Actually Shows

A histogram sorts continuous data into bins and stacks a bar for each one, with height representing either the raw count or the density of observations. Because the variable is continuous, the bars touch, unlike a bar chart, where gaps signal categorical data with no numeric order between groups.

That single visual gives you information a summary statistic cannot: modality (one peak, two peaks, or more), skewness (a long tail pulling right or left), gaps in the data, and unusual clusters near the tails. A histogram of exam scores might show two distinct peaks, one around 65 and one around 90, which a mean or median would flatten into a single, misleading middle value.

The catch is bin width. Choose too few bins and real structure disappears into a few fat bars. Choose too many and the chart turns into visual noise. Three common starting rules of thumb are Sturges’ formula, the square root rule, and the Freedman–Diaconis rule, which adjusts bin width based on the interquartile range and tends to hold up better on skewed data.

  • Sturges works well for small, roughly normal datasets.
  • The square root rule is a fast default for moderate sample sizes.
  • Freedman–Diaconis adapts to spread and outliers, making it a safer pick for messy real-world data.

Each rule arrives at a bin count differently. Sturges’ formula sets the bin count near ⌈log₂(n) + 1⌉, which assumes a roughly normal shape and doesn’t adjust for skew. The square root rule is blunter still — bin count ≈ √n — fast to compute but entirely agnostic to distribution shape. Freedman–Diaconis instead sizes each bin at roughly 2 × IQR × n^(−1/3), so a wider spread or a smaller sample produces wider bins, which is why it tends to hold up better once outliers or skew are present in the data.

Histograms also struggle with categorical variables (they need a bar chart instead) and with very small samples, where a handful of points can make bins look artificially smooth or artificially spiky. In those cases, a density plot, an ECDF, or a simple strip plot of raw points usually tells a more honest story.

Understanding Boxplots: The Five-Number Summary

A boxplot compresses a distribution into five numbers: the minimum, the first quartile (Q1), the median, the third quartile (Q3), and the maximum, with the box itself spanning Q1 to Q3, known as the interquartile range, or IQR. The line inside the box marks the median, not the mean, which is exactly why boxplots hold up so well against skewed data and extreme values.

The whiskers extend to the most extreme points that still fall within 1.5 times the IQR from the box edges, a convention known as the Tukey rule after statistician John Tukey, who introduced the box-and-whisker plot in his 1977 book Exploratory Data Analysis. Anything beyond that gets plotted as an individual dot and flagged as a potential outlier, worth a second look rather than automatic deletion. Some software distinguishes “mild” outliers, beyond 1.5×IQR, from “extreme” outliers, beyond 3×IQR — a useful distinction when you’re deciding how much scrutiny a flagged point deserves before you report it.

  • Strength: boxplots line up side by side cleanly, so comparing ten groups takes the same visual space as comparing two.
  • Strength: median and IQR are resistant to extreme values, unlike the mean and standard deviation.
  • Weakness: a boxplot cannot show you a bimodal distribution. Two peaks and one smooth hump can produce nearly identical boxes.
  • Weakness: with very small samples, the quartile estimates themselves get noisy, so the box can look tidy while hiding real instability.

Boxplots earn their keep in situations with many groups to compare or when a report needs to stay compact, such as comparing salary distributions across ten departments. They deserve more caution as a standalone tool when the underlying shape, not just the summary, is the actual question.

Histogram or Boxplot? A Side-by-Side Comparison

The decision usually comes down to what question you are actually answering, not which chart looks more polished.

Which chart to reach for, by task
Task Better chart Why
Exploring the shape of one variable Histogram It is the only one of the two that reveals modality directly.
Comparing several groups at once Boxplot Scales far better across five, ten, or twenty categories than a wall of histograms.
Hunting for outliers Boxplot The Tukey rule flags them automatically rather than requiring a visual guess.
Checking for multimodality before a test that assumes a single peak Histogram A boxplot will not warn you — two peaks and one hump can look identical.

Sample size matters more than most students expect. As a rule of thumb, aim for roughly 30 observations per group before trusting either chart’s summary as stable. Below that threshold, both histograms and boxplots can mislead, since bin heights and quartile positions swing wildly with just one or two added points. Practitioner guidance on small-sample visualization recommends plotting the raw points themselves, strip plots, dot plots, or beeswarm plots, once n drops much below 30, rather than dressing up noisy summary statistics.

There is also a trade-off in honesty versus tidiness. A histogram’s bin choice is subjective and needs disclosure; a boxplot’s quartiles are computed automatically but hide shape by design. When neither tool fits the job cleanly, violin plots, density plots, an ECDF, or small multiples of histograms across groups often split the difference better than forcing one chart to do two jobs. A violin plot in particular mirrors a density curve around a central axis, so it keeps the shape information a boxplot discards while still letting you line several groups up side by side — the trade-off is that it needs a reader who already knows how to read a density shape, which makes it a weaker choice for a mixed or non-technical audience than a boxplot.

How to Choose and Report: A Decision Guide

Start by confirming the variable type. Continuous data can go in either a histogram or boxplot; categorical data needs a bar chart, not a histogram, since bars for categories require gaps while histogram bars for continuous bins touch.

Next, check your sample size and group count. One group with enough data to explore shape points toward a histogram. Several groups needing quick comparison point toward a boxplot. When you genuinely cannot decide, make both — it costs little and often reveals something a single chart would have hidden. In practice, “make both” means building the histogram first so you know whether the shape is unimodal, skewed, or bimodal, then building the boxplot on the same variable and checking that the two tell a consistent story — a boxplot that looks perfectly symmetric next to a histogram with an obvious right tail is a signal to look closer at how the quartiles were computed, not to trust whichever chart happens to look cleaner.

Histogram vs boxplot decision path A decision tree first asks whether the variable is categorical, then whether multiple groups are being compared, then whether there are at least 30 observations, to land on a bar chart, boxplot, histogram, or raw-point plot. yes no yes no yes no Is the variablecategorical? Use a bar chart —categorical dataneeds gapsbetween bars Are you comparingmultiple groups? Use a boxplot forcompact medians,IQR, and outliers Are there atleast ~30observations? Use a histogramto explore shape,modality, andskewness Plot raw points —strip, dot, orbeeswarm —instead
Figure 1. A stop-and-decide path from variable type through group count and sample size to the right chart.

Record your settings as you go, since reproducibility depends on it:

Before you publish a distribution chart

  • Bin rule recorded Which rule you used — Sturges, square root, or Freedman–Diaconis — and how many bins resulted.
  • Sample size logged n for every group shown in the chart.
  • Outlier rule stated Which rule flagged points as unusual, typically the 1.5×IQR convention.
  • Median and IQR reported in the caption State the numbers explicitly rather than leaving readers to eyeball the chart.
  • Bimodality claim double-checked A second peak should survive across two or three different bin widths before you call a distribution bimodal.

Statohub’s Recommendation: Combine a Histogram With a Marginal Boxplot

A common classroom-tested approach is to pair the two visualizations rather than pick one. A histogram gives you shape; a boxplot stacked as a thin marginal strip along the same axis gives you the median, IQR, and outliers without eating extra space, a pattern that shows up often in practitioner workflows for aligning histograms and boxplots on a shared axis.

A simple build order works well for students:

  • Compute bin width using Freedman–Diaconis, then draw the histogram.
  • Calculate the five-number summary and add the boxplot as a narrow strip above or below, aligned to the same x-axis scale.
  • Annotate the median line and any flagged outliers directly on the combined figure.

Skip the combined display for very small samples or when comparing many groups at once. Small multiples or a violin plot per group usually reads more clearly in those cases.

What I Tell Students Who Ask “Which One Should I Just Use?”

Default to boxplots for grading and side-by-side comparisons; they force you to report median and IQR instead of eyeballing a shape. Default to histograms for exploratory work, where the question is “what does this distribution actually look like?” Always caption both with n and whichever bin or outlier rule you used. If you are still unsure which observation drove your conclusion, that is the sign to show both plots, not neither.

Practice What You Just Read

Reading about bin widths and quartiles only gets you so far. The fastest way to internalize the difference between a histogram and a boxplot is to build both from the same dataset and watch where they disagree. Statohub’s Learn hub walks through distribution concepts step by step, and the Applied Statistics section shows the same decisions playing out on real datasets rather than toy examples.

For the actual math, Statohub’s calculators handle the median, quartiles, and IQR you need to caption either chart correctly, and the descriptive statistics guide breaks down exactly what those numbers mean before you plot anything. If outliers are your main concern, the 1.5×IQR outlier guide walks through the Tukey rule in the same detail used above. Pull a dataset you already have, run both a histogram and a boxplot on it, and see which one actually answers your question.

Sources

Sources

  1. Atlassian — The Complete Guide to Histograms (bin-width rules) Atlassian
  2. NIST/SEMATECH e-Handbook of Statistical Methods — Box Plot National Institute of Standards and Technology
  3. NIST/SEMATECH e-Handbook — Detection of Outliers (1.5×IQR / Tukey rule) National Institute of Standards and Technology
  4. Krzywinski, M. & Altman, N. — "Points of Significance: Visualizing samples with box plots," Nature Methods (2014) Nature Methods
  5. John W. Tukey, Exploratory Data Analysis (Addison-Wesley, 1977) Addison-Wesley
  6. GeeksforGeeks — Merge and Perfectly Align Histogram and Boxplot Using ggplot2 in R GeeksforGeeks
  7. Mr. Slope Guy — Histogram vs Box Plot: When to Use Each for Describing Data Mr. Slope Guy

FAQ

Frequently asked questions

What is a disadvantage of using a boxplot instead of a histogram?
A boxplot hides the shape of the distribution, so a bimodal dataset with two distinct peaks can look identical to a smooth, single-peaked one once it is compressed into five numbers.
What do boxplots show that histograms don't?
Boxplots explicitly mark the median, the interquartile range, and flagged outliers using the 1.5×IQR rule, and they let you line up many groups side by side in a fraction of the space a row of histograms would need.
How do I match a histogram to a boxplot?
Align both plots to the same x-axis scale, then place the boxplot as a narrow marginal strip above or below the histogram so the box's quartiles and whiskers line up directly with the corresponding bars.
When should I use a bar chart instead of a boxplot or histogram?
Use a bar chart when your variable is categorical, since bar charts need gaps between bars to signal unordered categories, while histograms use touching bars specifically because the underlying variable is continuous.