Statohub Browse calculators
Machine Learning Statistics Practitioner guide

Avoid Data Leakage: 5 Step Train Test Split Checklist

A practical train test split guide for sklearn: choosing 80/20 vs 90/10 ratios, stratified splits, avoiding leakage, and when to use a holdout set instead.

By Statohub Editorial Team Published October 2026Reviewed October 202612 min read

A train-test split reserves part of your dataset, untouched, to measure how well a model generalizes to data it has never seen. The typical default is an 80/20 or 70/30 split, with the larger portion for training and the rest sealed away for a final, unbiased check. Get this step wrong and every metric you report afterward becomes suspect, no matter how sophisticated the model is.

Key takeaways

Point Details
Small test sets are noisy Fewer than 1,000 test samples gives an unreliable performance estimate — prefer cross-validation at that scale.
Respect structure in the data Grouped data (patients, users, stores) needs GroupKFold or GroupShuffleSplit; time-ordered data needs TimeSeriesSplit, or random splitting leaks information.
Split before you preprocess Fitting a scaler, encoder, or feature selector on the full dataset before splitting leaks test information and inflates reported performance.
Ratio depends on dataset size Large datasets (100,000+) can afford 90/10 or 95/5; medium datasets use 80/20 or 70/30; small datasets (under 1,000) are better served by k-fold cross-validation.

What Are Training, Validation, and Test Sets Used For?

Three datasets do three distinct jobs, and confusing them is where most beginners run into trouble. The training set teaches the model its parameters. The validation set helps you tune hyperparameters and compare candidate models without contaminating your final judgment. The test set, also called a holdout set when it has never been touched during development, exists solely to answer one question at the end: how will this model perform on data it has never encountered?

A simple two-way split (train and test only) works fine when you are not tuning any hyperparameters and just want a quick performance estimate. The moment you start comparing model architectures or adjusting settings like tree depth or regularization strength, you need a third partition, or you risk leaking test information into your decisions through repeated peeking.

Cross-validation offers a more efficient alternative to a fixed validation set, especially with limited data. According to Wikipedia’s overview of training, validation, and test data sets, a holdout set is specifically a test set that has never been used for any prior decision, and cross-validation or bootstrapping are recommended alternatives when sample size is small.

  • Training set: fits model parameters
  • Validation set: tunes hyperparameters and model choice
  • Test/holdout set: gives the final, one-time performance estimate

Common Split Ratios and How to Choose One

The 80/20 split is the most common default, with 70/30 a close second and more skewed splits reserved for very large datasets where even a small portion still contains thousands of examples. These ratios exist because they balance two competing needs: enough training data to learn real patterns, and enough test data to get a statistically stable performance estimate.

Typical test-set percentages by split ratio Bar chart of four common split ratios and the share of data held out for testing: 95/5 and 90/10 (large datasets) hold out 5% and 10%; 80/20 and 70/30 (medium datasets) hold out 20% and 30%. 0 9.75 19.5 29.25 39 5 95/5 (large) 10 90/10 (large) 20 80/20(medium) 30 70/30(medium) Split ratio Test set %
Figure 1. Typical test-set percentage by common split ratio, from a large-dataset 95/5 split up to a medium-dataset 70/30 split.

Dataset size should drive your decision more than habit. A dataset with 500,000 rows can afford a 90/10 split because 10% still gives you 50,000 test examples. A dataset with 300 rows cannot, and 20% for testing (60 rows) may already be too noisy to trust.

Rule of thumb: as your dataset shrinks, lean toward cross-validation instead of a single large holdout, since a small test set produces a performance estimate with wide, unstable variance.

  • Large datasets (100,000+): 90/10 or even 95/5 often works
  • Medium datasets (1,000 to 100,000): 80/20 or 70/30 is standard
  • Small datasets (under 1,000): prefer k-fold cross-validation over a fixed split

Research on model validation confirms that standard ratios of 70/30 and 80/20 are dataset-size and complexity dependent rather than fixed rules, and the same review recommends repeated cross-validation, with nested cross-validation for hyperparameter tuning, when sample sizes run small.

How to Perform a Train-Test Split in Python

The train_test_split function from sklearn.model_selection is the tool nearly everyone reaches for first, and for good reason: it handles shuffling, proportional sizing, and stratification in one line.

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,
    random_state=42,
    stratify=y
)

According to the scikit-learn documentation, this function is essentially a convenience wrapper around ShuffleSplit, and its behavior hinges on a handful of parameters worth understanding rather than memorizing:

train_test_split parameters and what each one controls
Parameter What it controls
test_size The proportion (or absolute count) of data held out for testing; 0.2 means 20%.
train_size An alternative way to specify the training proportion directly; rarely needed alongside test_size.
random_state A fixed integer seed so the split is reproducible every time you rerun the code.
shuffle True by default; randomizes row order before splitting. Must be set to False for ordered data like time series.
stratify Pass your target array (y) to preserve class proportions between train and test sets.

Skipping random_state means your split changes every run, which makes debugging and comparing model versions nearly impossible. Setting an absolute test size (test_size=1000 instead of 0.2) is sometimes clearer than a percentage when your dataset size might change over time.

Keeping Class Balance and Respecting Grouped Data

Stratified splitting prevents a subtle but damaging problem: if your target has a 90/10 class imbalance and you split randomly, you might end up with a test set that is 85/15 or 95/5 purely by chance, skewing every metric you calculate afterward. Passing stratify=y to train_test_split keeps the same class proportions in both subsets.

Grouping is a separate issue entirely, and stratification does not solve it. If your dataset contains multiple rows per patient, per user, or per store, a random split can put some of a person’s records in training and others in testing. The model then partly “memorizes” that person rather than genuinely generalizing, and your test score becomes inflated.

  • Use GroupKFold when you need k-fold cross-validation but must keep all rows from the same group together
  • Use GroupShuffleSplit when you want a single random split but still need groups kept intact
  • Watch for small groups: if one group has very few members, it may end up entirely in one subset, creating imbalance you cannot easily fix

Why Time-Series Data Needs a Different Splitting Approach

Random shuffling destroys the one thing time-series data depends on: order. If you shuffle a sales dataset before splitting, your model can end up training on next month’s data and testing on last month’s, which is a form of information leakage that makes results look far better than they will in production.

The fix is a forward, or “rolling window,” validation scheme where training data always precedes test data chronologically. Scikit-learn’s TimeSeriesSplit handles this directly, generating successive train/test splits that respect time order and supporting a gap parameter to leave a buffer between training and test windows, plus max_train_size to cap how much history each fold uses.

  • Never shuffle time-ordered data before splitting
  • Add a gap when your production model would have a delay between training and prediction
  • Match your test window length to how far ahead you actually need to forecast in practice

Statohub’s guide on why random cross-validation does not work for time-series data goes deeper into rolling-window mechanics if forecasting is your main use case.

How to Prevent Data Leakage When Splitting Data

Leakage happens whenever information from outside the training set influences the model, and it is the single most common reason a model performs beautifully in development and poorly in production. The three usual culprits are preprocessing on the full dataset before splitting (so your scaler “sees” test values), target leakage (a feature that encodes the outcome, like using a “cancellation date” column to predict cancellation), and selection leakage (running feature selection or hyperparameter search on the full dataset instead of only the training fold).

  1. Split your data first, always, before fitting any scaler, encoder, or imputer.
  2. Wrap preprocessing and modeling into a single sklearn.pipeline.Pipeline, so fit only ever touches training data and transform is applied consistently to the test set.
  3. Audit features for anything that could not have existed at prediction time in the real world.
  4. Fix a random_state across your whole workflow so splits and results are reproducible for review.

Large-scale research examining leakage across 2,047 benchmark datasets found that selection leakage and temporal contamination tend to produce the largest inflation in reported performance, more than other leakage types, which makes those two the priority to eliminate first.

Scikit-learn’s own guidance on common pitfalls echoes the same principle: split before you preprocess, every time, no exceptions.

A Step-by-Step Checklist for Choosing a Split Strategy

Working through these questions in order will get you to a defensible split strategy faster than guessing:

Choosing a split strategy

  1. Check independence Are your rows genuinely independent, or do some belong to the same user, patient, or store? If grouped, plan for GroupKFold or GroupShuffleSplit.
  2. Check for time order If timestamps matter, rule out random shuffling immediately and move to TimeSeriesSplit or a rolling-window approach.
  3. Assess dataset size Under roughly 1,000 rows, favor cross-validation over a single holdout; above that, a fixed split becomes more stable.
  4. Assess class balance If your target is imbalanced, stratify the split and consider whether cross-validation with stratified folds gives a steadier estimate.
  5. Decide on nested CV If you are both tuning hyperparameters and selecting between model types, use nested cross-validation so the outer loop still gives an honest, leakage-free performance estimate.

Checking independence first matters because it changes every downstream decision. A review of independence assumptions in machine learning notes that real-world data frequently violates the i.i.d. assumption that random splitting relies on, which is exactly why grouping and time structure need to be ruled out before you touch test_size.

Statohub Applied Example: Splitting a Small Clinical Dataset

Too few to trust.

The more defensible approach: use stratified 5-fold cross-validation across the full dataset for model comparison and hyperparameter tuning, then report the average performance across folds rather than a single train-test split score.

  • Confirm the target split ratio holds in every fold using StratifiedKFold
  • Fit any scaler or encoder inside a Pipeline so each fold’s training data is processed independently
  • If fold-to-fold accuracy or recall swings by more than a few percentage points, treat the estimate as unstable and collect more data before trusting it

Statohub’s Applied Statistics hub walks through similar small-sample workflows where cross-validation replaces a single holdout, and the PMC review on validation procedures backs the same recommendation for datasets this size.

What Beginners Get Wrong About Splitting Data

The most common mistake is not technical, it is behavioral: peeking at the test set repeatedly while tuning a model, then acting surprised when production performance disappoints. Every time you check test performance and adjust something in response, you have effectively used the test set for training, whether you fit the model on it or not.

The second mistake is preprocessing before splitting, which quietly leaks distributional information from test rows into your scalers and encoders. Document your split ratio, your random seed, and your preprocessing order every time, and treat your holdout set as something you touch exactly once, at the very end.

Statohub Resources to Try Next

Splitting data correctly is a skill you build through repetition, not memorization, and Statohub’s Machine Learning Statistics section covers the evaluation concepts that follow naturally after your split, including cross-validation design and confusion matrix interpretation. If you are still deciding how large your holdout needs to be, Statohub’s calculators can help you reason through sample-size questions before you commit code to a notebook.

For readers who want to see split-aware data workflows applied outside of pure machine learning, this walkthrough of building a small data pipeline shows similar discipline around keeping data stages separate and reproducible. And if the statistical reasoning behind sample size, variance, and estimation feels shaky, Statohub’s Learn hub builds those foundations in plain English before you ever touch a line of code. Start with whichever gap in your workflow is costing you the most confidence right now.

Sources

Sources

  1. scikit-learn — train_test_split documentation scikit-learn
  2. scikit-learn — Cross-validation: evaluating estimator performance scikit-learn
  3. PMC — Model validation procedures and data splitting (review) National Center for Biotechnology Information
  4. Which Leakage Types Matter? A Quantitative Landscape Across 2,047 Benchmark Datasets (arXiv, 2026) arXiv
  5. Wikipedia — Training, validation, and test data sets Wikipedia
  6. APXML — Common split ratios APXML

FAQ

Frequently asked questions

Why Do We Use an 80/20 Train-Test Split?
The 80/20 ratio balances two needs: giving the model enough data to learn genuine patterns while keeping enough held out to produce a stable performance estimate. It is a common default rather than a strict rule, and the PMC review on validation procedures notes the right ratio depends on dataset size and model complexity.
What Is a Good Train-Test Split Ratio?
For most medium-sized datasets, 80/20 or 70/30 works well. Larger datasets can shift toward 90/10 since even a small percentage still yields a sizable test set, while very small datasets are often better served by cross-validation instead of a single fixed split.
How Do I Split a Dataset Into Train and Test Sets?
In Python, the standard approach uses sklearn.model_selection.train_test_split, passing your features and target along with a test_size and a random_state for reproducibility. Add stratify=y when your target classes are imbalanced, and set shuffle=False if your data has a meaningful time order.
How Do I Import Train_test_split in Python?
Import it directly from scikit-learn with from sklearn.model_selection import train_test_split. It is part of the model_selection module alongside related tools like GroupKFold, GroupShuffleSplit, and TimeSeriesSplit, as documented in the scikit-learn reference.
Should I Use Cross-Validation Instead of a Single Train-Test Split?
Cross-validation gives a more stable performance estimate than a single split, especially with smaller datasets, because it averages results across multiple folds instead of relying on one holdout. A single train-test split is still useful and faster when your dataset is large enough that one held-out portion gives a reliable estimate on its own.