Optimizing an AI-Silicon Portfolio: Markowitz Mean-Variance with pypfopt, Done Right

Optimizing an AI-Silicon Portfolio: Markowitz Mean-Variance with pypfopt, Done Right

Ten tickers, one economic engine. If you build portfolios out of NVDA, AMD, AVGO, TSM, ASML, MU, INTC, QCOM, TXN, and ARM, you are not holding ten independent bets — you are holding one leveraged position on global AI compute demand, sliced ten ways. Equal-weighting hides that. Markowitz’s 1952 paper gives you the framework to see it; estimation error breaks naive implementations of it. This post walks through a working mean-variance pipeline with PyPortfolioOpt — shrinkage, bounds, and regularization included — and shows when to reach for something else.

The AI-Silicon Portfolio and Why It Needs Optimization

The ten names cover the semiconductor stack: GPUs and AI accelerators (NVDA, AMD), networking and custom ASICs (AVGO), foundry manufacturing (TSM), EUV lithography (ASML), HBM memory (MU), legacy CPUs and fabs (INTC), handsets and modems (QCOM), analog (TXN), and CPU IP licensing (ARM). Different businesses, same marginal buyer: hyperscaler AI capex.

That means the correlation structure, not just the volatilities, is the story. When a capex cycle turns, drawdowns arrive together across the whole panel, and an equal-weight portfolio — which looks diversified by ticker count — turns out to be a concentrated single-factor bet. Markowitz’s insight in Portfolio Selection was precisely that risk lives in the covariance matrix, so diversification is a quantitative optimization problem, not a counting exercise. This universe is the extreme case where that machinery earns its keep.

Markowitz Mean-Variance in One Page

The optimizer maximizes the trade-off between expected return and risk:

max_w  wᵀμ − λ·wᵀΣw    subject to    Σwᵢ = 1,  w_min ≤ w ≤ w_max

Here w is the weight vector, μ the vector of expected returns, Σ the covariance matrix, and λ the risk aversion parameter. Budget constraint plus bounds define the feasible set; the frontier is the locus of return-maximizing portfolios at each risk level.

The formulation is exact. The inputs are not. μ and Σ must be estimated from historical data, and every practical failure of mean-variance optimization traces back to that estimation step.

The Estimation-Error Problem (Why Naive MVO Fails)

Sample statistics are noisy. The covariance matrix of ten assets over three to five years of daily data is estimable but error-prone; expected returns are far worse — annualized mean returns carry enormous standard errors even with decades of data. An unconstrained optimizer doesn’t smooth over those errors, it exploits them: it loads maximum weight on whichever asset’s in-sample Sharpe ratio was flattered by noise. Michaud (1989) called this “error maximization” — the portfolios that look best on paper are precisely the ones whose inputs were most wrong.

The canonical empirical rebuke is DeMiguel, Garlappi and Uppal (2009), who found that 1/N beats optimized portfolios out-of-sample across 7 datasets — a devastating result for naive MVO, and not a claim that 1/N is optimal, only that estimation error overwhelmed the optimization benefit in their tests. The engineering response is threefold: shrink the covariance estimate, constrain the weights, and regularize the objective. All three ship in pypfopt.

pypfopt Setup and Data Pull

Install the stack and pull daily closes for the ten tickers:

pip install pypfopt cvxpy yfinance pandas
import yfinance as yf
from pypfopt import expected_returns, risk_models

TICKERS = ["NVDA", "AMD", "AVGO", "TSM", "ASML",
           "MU", "INTC", "QCOM", "TXN", "ARM"]

# 5 years of daily adjusted closes (yfinance docs: https://ranaroussi.github.io/yfinance/index.html)
raw = yf.download(TICKERS, start="2020-01-01", auto_adjust=True, progress=False)
prices = raw["Close"].dropna()  # note: ARM listed Sep 2023, so dropna() starts the panel there

mu = expected_returns.mean_historical_return(prices)        # annualized historical means
S  = risk_models.CovarianceShrinkage(prices).ledoit_wolf()  # Ledoit-Wolf shrunk covariance

print(mu.round(4))

The PyPortfolioOpt User Guide documents the full data-handling surface. One practical caveat: dropna() aligns to the shortest history, so ARM’s 2023 IPO shortens the panel — decide deliberately whether to accept that or pull a shorter common window.

Shrinking the Covariance Matrix: Ledoit-Wolf

The raw sample covariance is an unbiased but high-variance estimate. Ledoit and Wolf (2004) showed the fix: blend it with a highly structured target (constant-correlation is the default in pypfopt) with an intensity chosen automatically from the data. The shrunk estimate has higher bias but much lower estimation error — a strictly better input to the optimizer for correlated universes like this one, where sample off-diagonal entries swing wildly across windows.

The API is one line — risk_models.CovarianceShrinkage(prices).ledoit_wolf() — versus the raw estimate you’d get from risk_models.sample_cov(prices). Swap between them and compare the resulting weights; the difference is a live demonstration of the estimation-error problem.

Building the Efficient Frontier with pypfopt

With μ and S in hand, the MeanVariance module exposes the frontier. max_sharpe() finds the tangency portfolio; efficient_return() and efficient_risk() place you at any point on it. First, the naive version — deliberately unconstrained:

from pypfopt import EfficientFrontier

# "Before": default (0, 1) bounds — watch the corner behavior
ef_naive = EfficientFrontier(mu, S)
ef_naive.max_sharpe()
print(dict(ef_naive.clean_weights()))

Run this and you’ll typically see the weight mass piled into one or two of the highest recent performers rather than spread across the ten names — the error-maximization pathology in action. The full frontier can be traced by sweeping either method:

# Sweep: minimum-risk portfolio for a grid of return targets
for target in [0.15, 0.25, 0.35]:
    ef_t = EfficientFrontier(mu, S)
    ef_t.efficient_return(target_return=target)
    print(target, dict(ef_t.clean_weights()))

Taming Corner Solutions: weight_bounds and L2 Regularization

Two levers fix the corners. First, weight bounds — a floor of 2% forces diversification into every name, a cap of 35% limits single-name concentration. Second, L2 regularization, which adds a penalty on weight dispersion to the objective, shrinking all weights toward the equal-weighted center:

from pypfopt import objective_functions

# "After": efficient_risk with weight_bounds=(0.02, 0.35) set in the constructor,
# plus L2 regularization (gamma=1)
ef = EfficientFrontier(mu, S, weight_bounds=(0.02, 0.35))
ef.add_objective(objective_functions.L2_reg, gamma=1)
ef.efficient_risk(target_vol=0.30)   # maximize return s.t. annualized vol <= 30%

weights = ef.clean_weights()         # rounds dust weights to zero, enforces limits
print(weights)
print("Ret %.2f%% | Vol %.2f%% | Sharpe %.2f" % ef.portfolio_performance())

Before/after, the qualitative shift is consistent: the unregularized max-Sharpe vector concentrates; the bounded, L2-regularized efficient_risk vector spreads across all ten names while respecting your concentration cap. The exact weights depend on your data window — rerun it yourself rather than trusting any printed vector, including in tutorials. gamma is your tunable dial between Markowitz conviction and equal-weight humility: gamma=1 is a sensible default, higher values push harder toward 1/N.

Risk Measures Beyond Variance

Variance penalizes upside and downside symmetrically. If what you actually fear is tail drawdown, pypfopt offers EfficientCVaR, implementing Rockafellar and Uryasev (2000) — CVaR is the conditional expectation of loss beyond VaR, not VaR itself — optimized over scenario returns. If you distrust expected returns entirely, HRPOpt implements Lopez de Prado’s Hierarchical Risk Parity, which builds weights from correlation clusters without inverting Σ at all. Both are drop-in alternatives once your data pipeline is up.

Mean-Variance vs the Alternatives

Method Inputs needed Handles correlated AI-silicon risk? Main weakness When to use here
Mean-Variance (pypfopt, shrunk + bounds + L2) μ, Σ (both estimated) Yes, via covariance + explicit bounds/L2 Estimation error in μ; corners without regularization Default when you have reliable return estimates
HRP (Lopez de Prado) Returns only, no μ Yes, via hierarchical clustering — no Σ inversion No explicit return view; weights purely risk-based When you don’t trust μ at all
Black-Litterman Equilibrium weights + views Yes, via view-adjusted μ Requires articulating views; sensitive to view confidence When you hold strong NVDA-over-MU style views
Risk Parity Covariance only Partially — ignores expected returns entirely No return tilt; may underweight names with strong alpha As a naive baseline to sanity-check MVO
CVaR (Rockafellar-Uryasev) Scenario returns Yes — optimizes tail loss directly Needs many scenarios; LP not smooth When tail risk (drawdown) matters more than variance

FAQ

Do I need Ledoit-Wolf shrinkage if I already use weight bounds? Yes — they solve different problems. Bounds are constraints on the output; shrinkage repairs the quality of Σ as an input. Bounds can force diversification even with a noisy covariance matrix, but the optimizer still solves a distorted problem: misestimated covariances produce wrong risk contributions and wrong efficient_risk volatility targets. Both fixes are nearly free, so use both.

Why does max_sharpe sometimes return 100% in one ticker? Estimation error, amplified. The maximum-Sharpe solution is extremely sensitive to the sample inputs, and the optimizer awards the corner to whichever name’s in-sample numbers were flattered by noise — the exact mechanism behind the DeMiguel-Garlappi-Uppal out-of-sample result. The fix stack: weight_bounds=(0.02, 0.35), L2_reg with gamma=1, and clean_weights() to round dust.

Is pypfopt production-ready or just for research? It’s a thin, well-tested wrapper over cvxpy, which routes to mature convex solvers (Clarabel, ECOS, OSQP) — the same solver family used in production risk systems. The modeling layer is production-grade; what you’d add for a live system is your own execution, rebalancing schedule, and transaction-cost modeling.

What’s the difference between efficient_risk and efficient_return? efficient_risk(target_vol=0.30) maximizes expected return subject to annualized volatility at or below 30% — use it when your binding constraint is a risk budget. efficient_return(target_return=0.25) minimizes volatility subject to hitting a 25% return target — use it when the constraint is a return objective. They’re two parameterizations of the same frontier.

Can I plug my own return forecasts into pypfopt? Yes. mu is just a pandas Series indexed by ticker — replace the historical-mean estimate with your own vector and pass it straight into EfficientFrontier. For structured views, use pypfopt’s Black-Litterman module, which follows the original Black-Litterman (1992) framework to blend market-equilibrium returns with your absolute or relative views before optimization.

The Bottom Line

Desk-research verdict: pypfopt’s mean-variance optimization is the right first tool for this ten-ticker universe — if, and only if, you pair it with Ledoit-Wolf shrinkage, weight bounds around (0.02, 0.35), and L2 regularization. Strip those away and you reproduce the corner-solution pathology that Michaud (1989) diagnosed decades ago, with in-sample Sharpe ratios that evaporate out-of-sample — the failure mode DeMiguel, Garlappi and Uppal documented across seven datasets. Reach for HRP when you have no defensible μ and want purely risk-based weights; reach for Black-Litterman when you have strong, articulable views (say, NVDA acceleration versus MU pricing cycles) you want the optimizer to respect. Whatever you choose, validate on your own out-of-sample windows — no published result, and nothing in this post, substitutes for that test.

How This Guide Was Built

This is a desk-research piece: every quantitative claim is sourced from published academic work or official tool documentation, listed inline above — Markowitz (1952), Ledoit-Wolf (2004), Michaud (1989), DeMiguel-Garlappi-Uppal (2009), Rockafellar-Uryasev (2000), Lopez de Prado’s HRP paper, Black-Litterman (1992), and the PyPortfolioOpt, cvxpy, and yfinance documentation. No backtests were run in-house for this post, and no in-house performance numbers appear in it. The performance claims cited — including the 1/N out-of-sample comparison — come from the published papers referenced, not from our own runs. The code examples were written against the public pypfopt and yfinance APIs as documented at the linked doc pages and are runnable as shown; the exact weight vectors and portfolio performance figures they output will vary with your data pull date and window, so treat any single output as illustrative. Price data via yfinance uses Yahoo Finance’s unofficial API — for production work, use a licensed market-data feed. Nothing here constitutes investment advice or a claim of expected outperformance.

← Back to all posts