Fama-French Factor Regression in Python: OLS Without the Unit and Vintage Traps

A Fama-French factor regression is a few lines of statsmodels code. The hard part is handling the missing intercept, the percent-unit factor data, and the FIZ/CIZ vintage switch in Ken French’s data library. This tutorial shows the full workflow for quantitative traders and data scientists who already know pandas and regression basics.

E-E-A-T: Why trust this guide? This guide is written by the QuantBrainAI editorial team for readers who already work with pandas and regression. This tutorial is based on the original paper, the Ken French data library documentation and the statsmodels reference; it was not run against a live data feed for this article. Methodology follows the published Fama-French 1993 paper and the documented statsmodels API. File layouts, units, and vintage definitions change over time, so verify every detail against the current file you download before using results in research.

What the Fama-French model is (and why traders still use it)

The 1993 Fama-French model explains stock returns with three factors (market, size and book-to-market) and, in the same paper, two bond-market factors for maturity and default risk. Its stock factors remain the standard benchmark for separating manager skill from exposure to known risk premia. A later five-factor model adds profitability and investment factors and is a separate extension.

The classic model described in Common risk factors in the returns on stocks and bonds identifies five common risk factors. Three are stock-market factors: an overall market factor and factors related to firm size and book-to-market equity. Two are bond-market factors related to maturity and default risks.

For equity work, the three stock factors are the core:

  • MKT-RF: excess return of the market over the risk-free rate
  • SMB (Small Minus Big): return spread between small and large capitalization stocks
  • HML (High Minus Low): return spread between high and low book-to-market stocks

The three-factor version remains the standard first benchmark. If your strategy shows return, the first question is whether that return is alpha or simply loading on market, size, and value premia. A factor regression gives you that decomposition in a form risk managers, allocators, and researchers all understand.

You must supply your own portfolio return series for this tutorial. There are no sample returns provided here, because any invented series would obscure the data-alignment steps that matter in practice.

Getting the factor data from Ken French’s data library

Ken French’s library publishes Fama/French 3 Factors, Momentum, and 5 Factors as monthly, daily, and weekly TXT and CSV zips. Since January 2025 US returns use CRSP CIZ format with ex-date reinvestment, while legacy FIZ files ended December 2024. Never mix vintages in one regression sample.

The Ken French Data Library publishes Fama/French 3 Factors in TXT and CSV zip formats with monthly, daily, and weekly variants, alongside Momentum Factor (Mom) and Fama/French 5 Factors (2x3).

Two documentation facts are critical for replication:

  1. Vintage switch: Since the January 2025 data release the US research returns use CRSP Stock and Indexes Flat File Format 2.0 (CIZ). The legacy FIZ format was discontinued after the December 2024 release.
  2. Return definition change: In the legacy FIZ format monthly returns are month-to-month holding-period returns with dividends reinvested at month-end. In the CIZ format monthly returns are compounded daily returns with dividends reinvested on ex-dates.

Archived legacy data remains available in the 202412 archive. Any replication that mixes vintages is comparing different return definitions. Pick one vintage for your full sample period and document it. Do not splice FIZ and CIZ files to extend history.

import pandas as pd

def load_french_factors(path: str) -> pd.DataFrame:
    """
    Read the monthly block of a Ken French 3-factor CSV.
    Factors are in PERCENT per month. Returns month-end indexed frame.
    """
    rows = []
    with open(path, "r") as f:
        for line in f:
            parts = [p.strip() for p in line.strip().split(",")]
            if len(parts) >= 5 and len(parts[0]) == 6 and parts[0].isdigit():
                rows.append(parts[:5])
            elif rows:
                break  # monthly block ends at the first non-data line (annual block / footer)
    df = pd.DataFrame(rows, columns=["Date", "Mkt-RF", "SMB", "HML", "RF"])
    df[["Mkt-RF", "SMB", "HML", "RF"]] = df[["Mkt-RF", "SMB", "HML", "RF"]].astype(float)
    df.index = pd.to_datetime(df["Date"], format="%Y%m") + pd.offsets.MonthEnd(0)
    return df.drop(columns="Date").sort_index()

Always open the raw file once and confirm column names, date format, and header length. The library has changed formats before and may do so again.

Preparing your portfolio returns and the excess-return setup

Factor regressions require excess returns, calculated as portfolio return minus the RF risk-free rate from the French file. You must supply your own return series at matching frequency, aligned on identical dates. Clean timezone differences, period endpoints, and missing values before estimation to prevent silent misalignment and biased coefficients.

This tutorial uses monthly data. Daily data follows the same steps with the daily file, a YYYYMMDD date format, and no month-end normalization.

Steps:

  • Index your portfolio returns by date as a Series of monthly decimal returns (0.01 = 1%), e.g. portfolio_returns.
  • Use the RF column from the same factor file and vintage as your risk-free rate.
  • Inner-join on dates. No forward-fill of returns. Missing months should remain missing and be dropped pairwise.
  • Compute excess return as y = portfolio_return - RF after unit alignment.
def align_returns(
    portfolio_returns: pd.Series,
    factors: pd.DataFrame,
) -> pd.DataFrame:
    """
    portfolio_returns: monthly decimal returns indexed by date, reader-supplied.
    factors: output of load_french_factors() after to_decimal().
    Returns joined DataFrame with aligned month-end dates.
    """
    port = portfolio_returns.copy()
    port.index = pd.to_datetime(port.index) + pd.offsets.MonthEnd(0)
    port = port.sort_index()
    # Keep only overlapping dates, no filling
    joined = pd.DataFrame({"PORT": port}).join(factors, how="inner")
    joined = joined.dropna(subset=["PORT", "Mkt-RF", "SMB", "HML", "RF"])
    if joined.empty:
        raise ValueError("No overlapping dates. Check frequency and date index.")
    return joined

The loader and align_returns both map dates to month-end, so a month-start portfolio series joins correctly. Do not skip that normalization.

Running the OLS regression in statsmodels

statsmodels OLS does not include an intercept by default, so you must add one with add_constant to estimate alpha. Regress excess returns on MKT-RF, SMB, and HML with OLS(y, X).fit(), then inspect summary for coefficients, standard errors, t-statistics, and R-squared. Alpha is the intercept; slopes are factor loadings.

The statsmodels OLS class takes endog as 1-d and exog as nobs x k. An intercept is not included by default and must be added by the user. Pass missing='drop' to sm.OLS if your series contain NaNs.

import statsmodels.api as sm
from statsmodels.tools import add_constant
from statsmodels.regression.linear_model import RegressionResults

def run_ff_regression(
    portfolio_returns: pd.Series,
    factors: pd.DataFrame,
) -> RegressionResults:
    """
    Run 3-factor regression: (PORT - RF) ~ Mkt-RF + SMB + HML + intercept.
    portfolio_returns must be DECIMAL; factors are converted here.
    """
    joined = align_returns(portfolio_returns, to_decimal(factors))

    y = joined["PORT"] - joined["RF"]
    X = joined[["Mkt-RF", "SMB", "HML"]]
    X = add_constant(X, has_constant="add")
    validate_ff_inputs(joined, X)

    model = sm.OLS(y, X, missing="drop")
    results = model.fit()
    print(results.summary())
    return results

Call pattern for the reader:

# Reader supplies these two inputs:
# factors = load_french_factors("your_downloaded_factors.csv")
# my_returns = pd.Series(...)  # your monthly returns, indexed by date
# res = run_ff_regression(my_returns, factors)
# res.params, res.tvalues, res.rsquared hold coefficients, t-stats, R-squared

No output numbers are shown here by design. Your coefficients depend entirely on your portfolio series, sample period, and vintage. Inspect .params, .tvalues, .bse, and .rsquared on your own fitted object rather than comparing to a screenshot.

The two unit traps: percent returns and the intercept

French factors are published in percent, so decimal portfolio returns must be matched by dividing factors and RF by 100. Mixing units distorts all coefficients without error. The second trap is omitting the intercept, which forces alpha to zero. Both require explicit code guards, not visual inspection alone.

Per the Ken French Data Library documentation, factor returns are published in percent. Divide by 100 before comparing to decimal returns, and make that conversion explicit in code.

to_decimal() converts the factor columns explicitly. It does not guess units from magnitudes, so you state the units yourself.

def to_decimal(factors: pd.DataFrame) -> pd.DataFrame:
    """French files are in PERCENT. Return a copy in decimals (1.0 -> 0.01)."""
    out = factors.copy()
    cols = ["Mkt-RF", "SMB", "HML", "RF"]
    out[cols] = out[cols] / 100.0
    return out

The intercept trap is equally silent. Without add_constant, sm.OLS(y, X) fits a regression through the origin. There is no alpha term, slopes compensate, and R-squared is computed on a different basis. Always verify "const" in X.columns before calling .fit().

Reading and interpreting the output

Alpha is the intercept, measuring average return left unexplained by market, size, and value exposures. Slopes measure factor tilts: market beta, SMB loading, and HML loading. Use standard errors, t-statistics, and R-squared to judge precision and fit. This is statistical description of comovement, not evidence of future profitability or strategy quality.

Read the output as:

  • const / alpha: mean excess return not accounted for by the three factors over the sample. Positive alpha does not by itself imply a tradable edge — it reflects the model, period, and included factors.
  • MKT-RF loading: market beta. Near one suggests market-like comovement; far from one suggests net hedge or leverage.
  • SMB loading: size tilt. Positive implies small-cap tilt; negative implies large-cap tilt.
  • HML loading: value tilt. Positive implies value tilt; negative implies growth tilt.
  • Standard errors and t-stats: whether each loading is statistically distinguishable from zero in-sample.
  • R-squared: share of variance of excess returns accounted for by the factors in-sample.

Interpretation is statistical description, not a claim of profitability. A high R-squared means the factors track your portfolio, not that the portfolio is good. A significant alpha warrants further checks for omitted factors, transaction costs, and out-of-sample stability before any trading conclusion.

For extensions, the same library provides the 5-factor file and momentum. Add those columns to X only if your research question requires them, and keep frequency and vintage consistent.

Common mistakes and how to avoid them

Five common errors break Fama-French regressions: missing add_constant, percent-decimal mismatch, splicing FIZ with CIZ vintages, using raw instead of excess returns, and misaligned date indexes. Each is preventable with joins on dates, unit assertions, vintage checks, and explicit excess-return construction before calling OLS for estimation.

Checklist with guards:

  1. No intercept: Assert "const" in X.columns after add_constant. If missing, alpha is forced to zero.
  2. Percent vs decimal: Pass decimal portfolio returns and call to_decimal() once on the factors. Never regress decimal returns on percent factors.
  3. FIZ/CIZ splice: Use one vintage per regression. If your sample spans the January 2025 switch, choose either current CIZ history or the 202412 legacy archive for the full period. Document the choice.
  4. Raw instead of excess: Regress PORT - RF, not PORT. Using raw returns folds the risk-free rate into alpha.
  5. Date misalignment: Use how="inner" join and dropna(). Verify joined.index.min() and joined.index.max() match your intended sample. Watch month-start versus month-end timestamps and timezone-aware versus naive indexes.
def validate_ff_inputs(joined: pd.DataFrame, X: pd.DataFrame) -> None:
    assert "const" in X.columns, "Missing intercept: call add_constant before OLS."
    assert X[["Mkt-RF", "SMB", "HML"]].notna().all().all(), "NaNs in factors."
    # Vintage guard: record source file and date range in your notebook
    print(f"Sample: {joined.index.min().date()} to {joined.index.max().date()} n={len(joined)}")
    print("Confirm: single vintage (FIZ or CIZ), single frequency, excess returns used.")

Keep the workflow reproducible: store the downloaded filename, download date, vintage (FIZ/CIZ), and frequency alongside your portfolio definition. That provenance matters more than any single coefficient when revisiting research months later.

FAQ

Why is my alpha so large (or my betas near zero)?

A percent/decimal unit mismatch is the most common cause. If decimal returns (0.01) are regressed on percent factors (1.0), every slope is about 100 times too small. Call to_decimal() once, inspect raw CSV values, and confirm both sides use the same units before fitting.

Should I use the 3-factor or 5-factor model?

Depends on the question. The three-factor model is the classic benchmark and the focus of this tutorial — market, size, and value explain a large share of cross-sectional variation for many equity portfolios. If you are studying profitability or investment patterns, pull the 5-factor file from the same library and extend X. Keep interpretation consistent and avoid comparing loadings across models as if they are identical.

Do I need to worry about the CIZ change if my sample ends before 2025?

No, but never splice FIZ and CIZ vintages into one series. If your sample ends before 2025, you can use either vintage for that window, but do not append CIZ months to an FIZ history. Use the archive for pre-2025 data if you need to preserve the legacy definition, and note the vintage in your methods.

← Back to all posts