Purged K-Fold Cross-Validation Python: Stop Label Leakage

If you label each bar with a 20-bar forward return and then run standard k-fold cross-validation, adjacent bars share overlapping label windows, and shuffled folds put one bar’s test label into another fold’s training set. That is label leakage, and it inflates validation scores. Purged k-fold cross-validation fixes it by removing training observations whose labels overlap the test set, then applying an embargo for serial correlation.
How This Guide Was Built
This guide is based on the official documentation, the primary papers, and the vendors’ repositories — we did not run the code hands-on. All linked sources and version details were re-checked on the stated verification date. Every claim below traces to the Wiley book page, the scikit-learn and skfolio documentation, the purgedcv repository and paper, or the Bailey and López de Prado papers.
What Is Purged K-Fold Cross-Validation in Python?
Purged k-fold cross-validation is k-fold CV with two extra deletion steps applied to every training set: purging removes training observations whose label intervals overlap the test labels, and embargoing removes training observations that immediately follow the test set. The technique comes from Chapter 7 of Marcos López de Prado’s Advances in Financial Machine Learning (Wiley 2018), which you can find on the official Wiley page.
The formal definitions, quoted from the skfolio CombinatorialPurgedCV documentation, are: “Purging consists of removing from the training set all observations whose labels overlapped in time with those labels included in the testing set. Embargoing consists of removing from the training set all observations that immediately follow an observation in the testing set, since financial features often incorporate series that exhibit serial correlation (like ARMA processes).”
Why Does Standard K-Fold CV Leak on Financial Labels?
Standard k-fold assumes observations are i.i.d., and financial samples are not. When labels are computed over forward-looking windows — say a 20-bar forward return — bar i’s label contains information from bars i+1 through i+20. If fold boundaries split those bars between train and test, the model trains on data whose labels literally contain the test period’s outcome.
López de Prado devotes section 7.3 of the book to exactly this failure (“Why K-Fold CV Fails in Finance”) and section 7.5 to bugs in sklearn’s cross-validation for this domain, per the Wiley book page. The scikit-learn cross-validation user guide documents the i.i.d. assumption k-fold makes; it is a reasonable assumption for images or independent survey rows, and a broken one for overlapping financial labels.
The practical symptom is a validation score that looks strong out-of-sample but collapses in live trading, because the model effectively “saw” the answer during training. If your pipeline has this symptom, it is worth reading our guide on how to detect backtest overfitting in Python alongside this one.
What Does the Embargo Add to Purging?
The embargo handles leakage that purging cannot reach: information flowing forward from the test set into training observations that come after it. Purging removes training samples whose labels overlap the test window in either direction; embargoing additionally deletes a buffer of training observations immediately following the test set, because features built from serially correlated series (ARMA-style processes) carry traces of the test period.
Concretely, if your test fold ends at bar t and your features include rolling means or autoregressive lags, a training bar at t+1 may encode values observed during the test window. The embargo deletes that buffer. López de Prado ties the embargo length to the degree of serial correlation in the features, and it is applied after purging — both steps are described in the skfolio docs and Chapter 7 of the book.
When Does the Leak Bite Hardest?
The leak is worst when label horizons are long relative to fold size and labels are dense. A 20-bar forward-return label makes every training observation within 20 bars of the test boundary suspect; a 200-bar label makes it far worse. Triple-barrier labels, event-based samples, and any scheme where consecutive observations share outcome windows all fall into this category.
It also bites hardest on data with strong serial correlation — daily equity returns with volatility clustering, macro series, anything with regime persistence — because the embargo then matters as much as the purge. López de Prado’s GARP white paper on why most machine learning funds fail places data-related failures, including leakage, among the top reasons strategies die after deployment. If your labels are single-bar and your features are strictly backward-looking with no serial correlation, a fixed gap may suffice; otherwise, purge.
How Do You Code Split-Index Logic for Purging and Embargo in Python?
The core algorithm is a two-pass exclusion over label intervals: first drop training rows whose label interval overlaps the test fold’s interval, then drop rows whose labels end inside the embargo buffer after the test fold ends. You can implement this with boolean masks over integer bar indices in about twenty lines, and adapt it to your own label start and end arrays.
import numpy as np
def purged_embargo_train_idx(label_start, label_end, test_idx, embargo_size):
"""label_start/label_end: arrays of each sample's label interval.
test_idx: integer positions of the current test fold.
embargo_size: number of bars to embargo after the test fold's end."""
test_start = int(label_start[test_idx].min())
test_end = int(label_end[test_idx].max())
train_mask = np.ones(len(label_start), dtype=bool)
train_mask[test_idx] = False # exclude the test fold
# Pass 1 — purge: any training label overlapping the test label interval.
overlap = (label_start <= test_end) & (label_end >= test_start)
train_mask &= ~overlap
# Pass 2 — embargo: labels ending within the buffer after test_end.
embargo_zone = (label_end > test_end) & (label_end <= test_end + embargo_size)
train_mask &= ~embargo_zone
return np.where(train_mask)[0]
# Example: 20-bar forward-return labels, 5-bar embargo.
n = 1000
label_start = np.arange(n)
label_end = np.minimum(label_start + 20, n - 1)
train_idx = purged_embargo_train_idx(label_start, label_end,
test_idx=np.arange(600, 700),
embargo_size=5)
Pass the purged_size and embargo_size semantics you choose here consistently across folds. This snippet only builds indices; it asserts nothing about model performance.
Scikit-Learn TimeSeriesSplit gap vs mlfinlab vs skfolio vs purgedcv: Which Should You Use?
Your choice in 2026 comes down to license, maintenance, and how much leakage protection you need. Scikit-learn’s TimeSeriesSplit offers only a fixed sample-count gap; skfolio ships a maintained BSD-3-Clause CPCV splitter; purgedcv is a small MIT package covering the full López de Prado toolkit; mlfinlab is the canonical implementation but is now paid and closed.
| Tool | Leak protection | License / cost | Maintenance signal |
|---|---|---|---|
sklearn TimeSeriesSplit |
Fixed gap of samples before each test set (docs) |
Free, BSD | Part of scikit-learn 1.9.1 |
| mlfinlab (Hudson & Thames) | Canonical Chapter 7 purging/embargo (pricing) | “£100 (+VAT) per month, per user” + commercial license; repo license NOASSERTION | Last push 2023-10-02 (GitHub) |
skfolio CombinatorialPurgedCV |
Purging + embargo, combinatorial test paths (docs) | Free, BSD-3-Clause | v1.3.0, Python >=3.10, 2,431 stars, last push 2026-09-22 (GitHub) |
| purgedcv | Purging; time-, observation-, or fraction-based embargo; group-purged k-fold; CPCV; PBO/DSR (PyPI) | Free, MIT | v0.1.6, Python >=3.10, 35 stars, last push 2026-09-04 (GitHub) |
The critical distinction: sklearn’s gap parameter — added in version 0.24 and defined as the “Number of samples to exclude from the end of each train set before the test set” — is a fixed row count. It inspects no label intervals, performs no label-overlap purging, supports no fractional embargo, and has no group awareness. If your label horizon equals a constant number of bars and you compute the gap yourself, TimeSeriesSplit(gap=N) can be acceptable. The moment labels vary in length, or you need group or fractional embargo semantics, move to skfolio or purgedcv. mlfinlab remains the reference implementation historically, but per the purgedcv paper it “was relicensed as a paid closed-source product and cannot be a dependency for an open project,” and its repo has been stale since October 2023.
How to Drop Purged CV into a sklearn Pipeline with purgedcv
Purgedcv implements every splitter against the scikit-learn splitter protocol, so its objects drop directly into cross_val_score, GridSearchCV, and Pipeline without wrappers. Per the package paper, it is an “open, MIT-licensed, maintained, sklearn-native implementation checked against the original papers,” ships py.typed, passes mypy --strict, and is “pinned by 354 tests … at 98% line coverage.”
from sklearn.model_selection import cross_val_score, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import Ridge
from purgedcv import PurgedKFold # sklearn splitter protocol
# purgedcv takes the label intervals explicitly: prediction_times and
# evaluation_times are the start/end of each row's label horizon.
splitter = PurgedKFold(
n_splits=5,
prediction_times=label_start,
evaluation_times=label_end,
purge_horizon="5D", # purge labels overlapping the test span
embargo_observations=5, # trailing buffer after each test fold
)
pipe = Pipeline([
("scale", StandardScaler()),
("model", Ridge()),
])
# Works with any sklearn utility that accepts a cv splitter:
grid = GridSearchCV(pipe, param_grid={"model__alpha": [0.1, 1.0, 10.0]},
cv=splitter)
# grid.fit(X, y) # run on your own labelled data
scores = cross_val_score(pipe, X, y, cv=splitter) # X, y from your pipeline
Check the purgedcv documentation for the exact class names and constructor signatures in the version you install, since the package exposes several splitters (purged k-fold, group-purged k-fold, and walk-forward variants).
How Does skfolio’s CombinatorialPurgedCV Recombine Test Paths?
Skfolio’s CombinatorialPurgedCV goes beyond single-path validation: instead of producing one train/test sequence, it builds many combinations of test folds and recombines them into multiple backtest paths, each purged and embargoed. Where KFold “can recombine one single testing path,” CombinatorialPurgedCV “can recombine multiple testing paths from the combinations of the train/test sets,” according to the skfolio docs.
from skfolio.model_selection import CombinatorialPurgedCV
cv = CombinatorialPurgedCV(n_folds=10, n_test_folds=8,
purged_size=0, embargo_size=0)
for train_idx, test_idx in cv.split(X):
# train_idx / test_idx are position arrays for one combination
print(len(train_idx), len(test_idx)) # index sizes only — no scores
# Or hand it directly to a sklearn-compatible estimator:
# cross_val_score(model, X, y, cv=cv)
Set purged_size and embargo_size to your label horizon and serial-correlation buffer rather than leaving them at zero. Single-path evaluation is the traditional alternative — see our walk-forward analysis methodology — but combinatorial paths give you a distribution of paths to evaluate instead of one history, which is what the overfitting statistics in the next section consume.
Why Pair Purged CV with the Deflated Sharpe Ratio?
Purging and embargo remove leakage from the validation folds; they do nothing about selection bias across many strategy trials. If you test hundreds of configurations, the best cross-validated Sharpe ratio is still inflated by luck. That is why López de Prado pairs purged CV with the Deflated Sharpe Ratio (Bailey & López de Prado, JPM 2014) and the Probability of Backtest Overfitting (Bailey, Borwein, López de Prado, Zhu).
The purgedcv package ships the Probabilistic, Deflated, and Minimum-Track-Record Sharpe statistics alongside its splitters (PyPI), so the full Chapter 7-plus-validation workflow lives in one sklearn-native dependency. CPCV’s multiple recombined paths are exactly what PBO needs as input: more paths give a better estimate of how often the in-sample-best configuration underperforms out of sample. Treat leakage control and overfitting statistics as two layers of the same defense — our backtest overfitting detection guide covers the second layer in depth.
FAQ
Can I just use sklearn’s TimeSeriesSplit with a large gap?
Sometimes. If every label spans exactly N bars and your features are backward-looking, TimeSeriesSplit(gap=N) excludes a fixed buffer before each test set and may be adequate — you must compute N from your forecast horizon yourself. It gives no label-overlap purging, no fractional embargo, and no group awareness, so variable-length labels or grouped samples demand a purged splitter.
Does purged k-fold CV eliminate backtest overfitting?
No. Purging and embargo remove leakage from overlapping labels and serial correlation within the folds, but they do not correct selection bias across many trials, non-normal return distributions, or multiple-testing inflation. Pair them with the Deflated Sharpe Ratio and Probability of Backtest Overfitting to quantify how much of your result survives those effects.
Why is mlfinlab not the default recommendation anymore?
It is the canonical Chapter 7 implementation, but Hudson & Thames relicensed it as a paid closed-source product at “£100 (+VAT) per month, per user” plus a separate commercial license, and its public repo shows a last push of 2023-10-02 with a NOASSERTION license. For an open project, skfolio (BSD-3-Clause) or purgedcv (MIT) are maintained, free alternatives.
← Back to all posts

