Triple-Barrier Labeling and Meta-Labeling in Python: A 2026 Implementation Guide

Fixed-horizon labels judge a trade by where price finishes, so identical endpoint returns can hide opposite paths through stops, targets, and drawdowns. In 2026, triple-barrier labeling preserves that path, while meta-labeling filters a primary side call into an explicit precision-and-coverage decision. The implementation below connects both layers to purged, embargoed evaluation rather than treating label construction as evidence of an edge.
How This Guide Was Built
This guide was built from primary desk research represented by the Wiley AFML book plus the reproducible results of the orchestrator’s reference run, a synthetic GBM experiment with seed 42 and n_events 2479, without hands-on trading or vendor evaluation of any product or service.
The orchestrator’s reference run used 2,500 business days, daily sigma 1.2 percent, annualised vol 19.0 percent, start 2016-01-04, and end 2025-08-01; the environment used numpy 2.4.6 and pandas 3.0.5. The orchestrator genuinely executed this deterministic synthetic run for the guide, and no vendor product was tested. Scope: this guide is built from the primary documentation and papers plus a reproducible reference implementation run on a deterministic synthetic price path; no live market data, no brokerage and no capital were involved.
Why Fixed-Horizon Labels Fail
Fixed-horizon return-sign labeling fails because it ignores the path taken to the horizon, a limitation specified in Chapter 3 of Wiley AFML; it cannot encode stops, targets, or drawdown rules that determine whether a trade survived before its endpoint return is observed.
At a fixed horizon, a long trade that touches its stop and later recovers to a positive terminal return receives the same directional label as one that reaches its target without the same adverse excursion. A short mirrors that example. The label discards which constraint occurred first, so it cannot represent the actual execution rule. It can also make repeated paths look like abundant evidence unless event overlap is controlled.
Triple-Barrier Labeling in Python
Triple-barrier labeling records whether the upper, lower, or vertical barrier is touched first, following the event definition used in the arXiv 2504.02249 Korean-market study: upper is +1, lower is -1, and expiry without a touch is 0 for every entry event.
The function expects a monotonic close series and an events frame indexed by entry time, with a numeric side of 1 or -1. The construction follows Wiley AFML conventions, while this MLFinLab labeling implementation provides a direct implementation reference. The function scans only the forward path and uses np.argmax on two touch masks to select the earliest barrier.
import numpy as np
import pandas as pd
def triple_barrier_label(
close, events, pt_mult, sl_mult, vol_span=100, hold=20
):
close = close.astype("float64").sort_index()
events = events.sort_index()
if not close.index.is_monotonic_increasing or not close.index.is_unique:
raise ValueError("close must be sorted and have unique timestamps")
if not events.index.is_unique:
raise ValueError("events must have unique timestamps")
if pt_mult <= 0 or sl_mult <= 0 or hold < 1:
raise ValueError("multipliers must be positive and hold at least one bar")
if not events.index.isin(close.index).all():
raise ValueError("every event must exist in close")
returns = close.pct_change(fill_method=None)
daily_vol = returns.ewm(span=100).std().shift(1)
if vol_span != 100:
daily_vol = returns.ewm(span=vol_span).std().shift(1)
labels = pd.Series(index=events.index, dtype="int64")
first_touch = pd.Series(index=events.index, dtype="int64")
t1 = pd.Series(index=events.index, dtype=events.index.dtype)
for timestamp, event in events.iterrows():
side = float(event["side"])
if side not in (-1.0, 1.0):
raise ValueError("side must be 1 or -1")
sigma = float(daily_vol.loc[timestamp])
if not np.isfinite(sigma) or sigma <= 0:
raise ValueError("positive shifted volatility is required")
start_loc = close.index.get_loc(timestamp)
if not isinstance(start_loc, (int, np.integer)):
raise ValueError("close timestamps must be unique")
end_loc = min(start_loc + int(hold), len(close) - 1)
path = close.iloc[start_loc:end_loc + 1]
anchor = float(close.iloc[start_loc])
upper_price = anchor + side * pt_mult * sigma
lower_price = anchor - side * sl_mult * sigma
path_values = path.to_numpy()
touches = np.column_stack(
[
path_values >= upper_price,
path_values <= lower_price,
]
)
hit_rows = np.flatnonzero(touches.any(axis=1))
if hit_rows.size:
hit = int(hit_rows[0])
barrier = int(np.argmax(touches[hit]))
labels.loc[timestamp] = 1 if barrier == 0 else -1
first_touch.loc[timestamp] = barrier
else:
labels.loc[timestamp] = 0
first_touch.loc[timestamp] = 2
t1.loc[timestamp] = path.index[min(int(hit_rows[0]), len(path) - 1)] if hit_rows.size else path.index[-1]
return pd.DataFrame({"label": labels, "t1": t1})
The output stores label and t1 for each event. Close-only data resolve the first observed touch deterministically; OHLC data need an explicit policy when both barriers appear inside one bar. Never relabel an event after seeing which barrier was eventually reached. That would use future path information twice and turn the target itself into a look-ahead signal.
Volatility-Scaled Barrier Widths Without Look-Ahead
Barrier widths are volatility multiples based on the exponentially weighted estimator documented by pandas EWM, shifted by one bar to enforce look-ahead safety; bar t then cannot use its own return to set its barrier or influence its own event risk.
Use daily_vol = returns.ewm(span=100).std().shift(1) for the guide’s baseline. For a long event, the upper distance is pt_mult * sigma; for a short, the sign reverses the barrier direction. The lower width can use a separate multiplier when payoff and risk are asymmetric.
Longer volatility spans react slowly to abrupt regime changes, while shorter spans react quickly but can produce noisy widths. Compare spans only inside training folds. The pandas rolling documentation covers alternative estimators, while the QuantResearch lectures provide event-based labeling context.
Overlap, Concurrency, and Sample Uniqueness
Overlapping label spans create concurrency and reduce statistical independence, as Chapter 4 of Wiley AFML defines through average uniqueness; the orchestrator’s reference run reported avg_concurrency 7.12, avg_uniqueness 0.143, and effective_n 354.5 from n_events 2479 rather than treating every row as independent evidence.
An event’s uniqueness is the reciprocal of its mean concurrency over the label span. The reported average is therefore a diagnostic for dependence, while effective_n is a uniqueness-based estimate rather than a literal fractional trade. Use it as a diagnostic and possible bootstrap weight, not as a substitute for leakage control. Before model selection, apply purged K-fold cross-validation in Python so training events whose spans intersect a test interval are removed.
Meta-Labeling: A Second Model That Filters the First
Meta-labeling is a two-model architecture in which a primary model chooses the side and a secondary binary classifier decides whether to act, as Chapter 3.6 of Wiley AFML specifies; correctness becomes a tunable precision-versus-coverage filter for each primary event itself.
The primary side remains authoritative. A long primary call with a +1 triple-barrier label is meta-positive, as is a short call with a -1 label. A vertical 0 is not a correct directional call. This construction preserves the Hudson Thames meta-labeling study’s separation between proposing a side and deciding whether that side should be taken.
import numpy as np
import pandas as pd
def build_meta_labels(primary_side, triple_barrier_label):
if not primary_side.index.equals(triple_barrier_label.index):
raise ValueError("both inputs must use the same event index")
side = primary_side.to_numpy(dtype=float)
barrier = triple_barrier_label.to_numpy(dtype=float)
valid = (
np.isfinite(side)
& np.isfinite(barrier)
& np.isin(side, [-1.0, 1.0])
& np.isin(barrier, [-1.0, 0.0, 1.0])
)
if not valid.all():
raise ValueError("inputs contain invalid side or barrier labels")
y_meta = (np.sign(primary_side) == np.sign(triple_barrier_label)).astype(int)
return y_meta.rename("y_meta")
Train the secondary classifier on a chronological training subset and use its probability to reject, retain, or scale the primary event. The decision threshold is a research parameter rather than an automatic optimum. Meta-labeling can improve precision among accepted events, but it cannot create an expected edge when the primary signal contains no predictive information.
Purged K-Fold Cross-Validation with Embargo
Purged K-fold cross-validation with embargo removes training observations whose label intervals intersect a test fold, matching the interval purging implemented in the maintained purged-cross-validation package; plain time ordering alone does not remove that overlap from the training side of each fold.
Each event needs inclusive bar_start and bar_end values. The function forms chronological test blocks, purges intervals intersecting any test event, and removes entries occurring during the configured embargo after the latest test endpoint. A walk-forward analysis methodology can organize the outer research process. TimeSeriesSplit preserves order but does not know event intervals.
import numpy as np
def interval_overlap_mask(a_start, a_end, b_start, b_end):
return (a_start <= b_end) & (a_end >= b_start)
def purged_embargo_kfold(bar_start, bar_end, n_splits=5, embargo=1):
start = np.asarray(bar_start, dtype=np.int64)
end = np.asarray(bar_end, dtype=np.int64)
if start.ndim != 1 or end.shape != start.shape:
raise ValueError("bar_start and bar_end must be matching vectors")
if np.any(end < start):
raise ValueError("each event must end on or after its start")
if n_splits < 2 or embargo < 0:
raise ValueError("n_splits must be at least two and embargo nonnegative")
order = np.argsort(start, kind="stable")
folds = np.array_split(order, n_splits)
for test_idx in folds:
overlap = np.zeros(start.size, dtype=bool)
for test_event in test_idx:
overlap |= interval_overlap_mask(
start, end, start[test_event], end[test_event]
)
test_finish = int(end[test_idx].max())
future = order[start[order] > test_finish]
embargo_idx = future[start[future] <= test_finish + embargo]
train_mask = np.ones(start.size, dtype=bool)
train_mask[test_idx] = False
train_mask[overlap] = False
train_mask[embargo_idx] = False
train_idx = np.flatnonzero(train_mask)
if train_idx.size == 0:
raise ValueError("purging removed the entire training fold")
yield train_idx, test_idx
Purging addresses direct interval intersection; embargo adds a buffer after test information has ended. Both are necessary controls when labels remain open across multiple bars. Store the event intervals with every dataset split so this logic remains auditable.
A Numpy-Only Logistic Regression for Meta-Labeling
A NumPy-only gradient-descent logistic regression is sufficient to implement the secondary layer, following the basic optimization form documented for scikit-learn logistic regression; the orchestrator’s reference run used 4,000 iterations, learning rate 0.1, and L2 1e-3 across four z-scored features in total.
Z-score each feature using training-set statistics only, then pass the resulting matrix to the optimizer. In the orchestrator’s reference run, the exact specification was numpy logistic regression (GD, 4000 iters, lr 0.1, L2 1e-3), 4 features, z-scored, 70/30 chronological split. The deterministic optimizer consumes no random draws; the NumPy Generator documentation becomes relevant when adding bootstrap or randomized extensions.
import numpy as np
# sigmoid(z) = 1 / (1 + np.exp(-z))
def sigmoid(z):
z = np.clip(z, -35.0, 35.0)
return 1.0 / (1.0 + np.exp(-z))
def logistic_regression_gd(X, y, lr=0.1, l2=1e-3, iters=4000):
X = np.asarray(X, dtype=float)
y = np.asarray(y, dtype=float).reshape(-1)
if X.ndim != 2 or X.shape[0] != y.size:
raise ValueError("X must be two-dimensional and align with y")
design = np.column_stack([np.ones(X.shape[0]), X])
weights = np.zeros(design.shape[1], dtype=float)
for _ in range(iters):
probabilities = sigmoid(design @ weights)
gradient = design.T @ (probabilities - y) / y.size
gradient[1:] += l2 * weights[1:]
weights -= lr * gradient
return weights
The first returned weight is the intercept; the remaining weights correspond to standardized features. Keep the chronological split outside the fitting function, evaluate probabilities on held-out events, and inspect precision, recall, F1, AUC, and coverage together.
What the 2026 Reference Run Shows
The orchestrator’s reference run shows that barrier choice changes label distributions and holding periods on a deterministic synthetic GBM path, but it does not establish predictive edge; classification metrics such as AUC must be evaluated separately, as described in scikit-learn model evaluation.
Barrier sweep from the orchestrator’s reference run.
| Barrier configuration | Upper touched % | Lower touched % | Vertical barrier % | Mean holding bars |
|---|---|---|---|---|
| pt1.0_sl1.0_hold20 | 50.7 | 49.2 | 0.1 | 2.77 |
| pt2.0_sl2.0_hold20 | 47.8 | 50.2 | 1.9 | 6.12 |
| pt3.0_sl3.0_hold20 | 42.3 | 43.7 | 14.0 | 7.65 |
| pt2.0_sl1.0_hold20 | 38.0 | 61.5 | 0.5 | 4.15 |
| pt1.0_sl2.0_hold20 | 61.4 | 38.3 | 0.4 | 4.21 |
| pt2.0_sl2.0_hold5 | 26.3 | 26.9 | 46.8 | 1.65 |
| pt2.0_sl2.0_hold60 | 48.9 | 51.0 | 0.1 | 6.55 |
For pt2.0_sl2.0_hold20, the orchestrator’s reference run recorded 1245 lower touches, 1186 upper touches, and 48 vertical touches, representing 50.2 percent, 47.8 percent, and 1.9 percent of n_events 2479. In the orchestrator’s reference run, the five-bar configuration produced 46.8 percent vertical outcomes and 1.65 mean holding bars, while the 60-bar configuration produced 0.1 percent and 6.55 bars. The vertical cap is therefore a maximum holding horizon, not a fixed realized duration.
The MLFinLab GitHub repository, its labeling source, and a triple-barrier meta-labeling example clarify implementation mechanics but do not validate a strategy. The precision score documentation also shows why accepted-case counts and denominators must accompany a precision value.
The orchestrator’s reference run assigned n_primary_signals 1171 to the meta-labeling problem, with n_train 819 and n_test 352; base_rate_train was 0.457 and base_rate_test was 0.491. It reported auc_test 0.587. At threshold 0.5, precision was 0.561, recall was 0.428, F1 was 0.485, and threshold 0.5 coverage was 0.375.
At threshold 0.6, the orchestrator’s reference run reported precision 0.75, recall 0.017, and F1 0.034, but threshold 0.6 coverage was 0.011 because only 4 of 352 test signals were accepted. Under ROC AUC scoring, the modest test result on deterministic synthetic prices is not evidence of a tradeable edge.
Pitfalls That Corrupt Triple-Barrier Labels
Seven failure modes corrupt triple-barrier and meta-labeling pipelines: barrier leakage, volatility look-ahead, label overlap, class imbalance, calibration leakage, threshold miscalibration, and mistaking label balance for edge, consistent with AFML and cross-validation guidance across the complete research and evaluation workflow as a whole.
Control each failure at its source: freeze event information before labeling, shift volatility, purge overlapping intervals, inspect labels by fold, select multipliers in training, and report threshold coverage. A nearly balanced label histogram says only that outcomes were common, not that they were predictable. Use model evaluation across metrics and backtest overfitting detection in Python before accepting a pipeline.
The Bottom Line
The bottom line is that triple-barrier labeling improves path alignment but creates no edge by itself; the orchestrator’s reference run produced auc_test 0.587 and threshold 0.5 precision 0.561, which do not establish a tradeable signal under ROC AUC evaluation for this synthetic sample.
Use triple-barrier labels for path-aware supervision, retain the primary side model, and use meta-labeling as a probability filter. Evaluate through purged and embargoed splits, then model transaction costs in a backtest before sizing. Precision, coverage, effective sample count, turnover, and costs belong in the same decision record.
FAQ
This FAQ answers common implementation questions about triple-barrier labeling, meta-labeling, split design, and interpretation of the 2026 synthetic reference run, using the same Wiley AFML principles and the orchestrator’s measured results rather than claims about live trading or financial performance.
Do I need a profitable primary model before meta-labeling?
A primary model only needs side calls with enough recall for the secondary classifier to filter. Meta-labeling evaluates whether each proposed direction matches the subsequent triple-barrier outcome; it cannot manufacture predictive content when the primary signal has none. A rich event set can still produce a useless filter if every direction is effectively random.
How should I choose the triple-barrier multipliers and vertical hold period?
Choose multiplier and hold-bar values on training folds only, using the reference sweep pattern as a sanity check rather than a universal optimum. Wider barriers generally extend holding periods and increase vertical-expiry share, while asymmetric barriers skew upper versus lower touches. Finalize choices before test-fold evaluation to avoid calibration leakage.
Why use a binary meta-label instead of directly predicting the triple-barrier outcome?
The binary target answers whether the primary side call was correct, allowing the secondary probability to tune precision and coverage without changing the side model. Directly predicting the original triple-barrier label mixes direction and timing into one multiclass outcome, weakening the clean separation between proposing a trade and deciding whether to take it.
Can I substitute scikit-learn TimeSeriesSplit for the purged/embargoed split?
Plain TimeSeriesSplit preserves chronological order, but it still allows label information from overlapping spans to cross boundaries; purged K-fold cross-validation in Python handles that problem by removing intersecting training intervals and embargoing observations immediately after each test fold when labels span bars.
Does the 2026 reference run demonstrate a tradeable edge?
The orchestrator’s reference run produced auc_test 0.587 and threshold 0.6 coverage 0.011, so the apparently stronger precision at 0.6 applies to only 4 of 352 test signals. That selective coverage, combined with a modest test AUC on deterministic synthetic prices, is not evidence of a tradeable edge or performance that should be expected in markets.


