← All articles

Does the VIX predict crypto risk? Not in any single year.

Short answer: the headline number says yes, and it is lying to you. Scored out-of-sample against a forward drawdown label on SOL, the VIX comes out at an AUC of 0.537 — where 0.50 = coin flip and 1.00 = perfect. That looks like a small but real edge. Then you split the same data by calendar year and the average drops to 0.441, with only three of seven years above a coin flip. The pooled score is higher than the best single year on record. This post is about how that happens, because it is one of the easiest ways to fool yourself with a backtest.

Why anyone asks

The VIX is the options market's estimate of how much the S&P 500 will swing over the next month — Wall Street's fear gauge. Crypto traders watch it because the intuition is compelling: when equities get scared, risk assets sell off together, and crypto is the highest-beta risk asset in the room. "Watch the VIX" is standard advice.

It is also testable, which is the entire point of this series. We already did this to crypto's own sentiment gauge and found the Fear & Greed Index scores a coin flip over six years. The VIX is the more serious candidate, and it fails in a more interesting way.

How we scored it

Identical method to the last one, so the two are comparable:

The result

Pooled across everything: AUC 0.537. Modest, positive, the kind of number you would happily put in a deck.

By year:

Average of those: 0.441. Best year: 0.528. The pooled score beats every individual year in the sample. That is not a rounding artefact or a bad year dragging things down — it is structural, and once you see the mechanism you will spot it everywhere.

How pooling manufactures skill

AUC asks a pairwise question: take one risky day and one calm day, and how often does the indicator rank them correctly? Pool six years together and most of those pairs are cross-year pairs — a day from 2022 against a day from 2021.

Now think about what the VIX was doing. 2022 was a high-VIX year and a genuinely dangerous one for crypto. 2021 was a lower-VIX year and, for most of it, calmer. So the pooled test keeps asking "is this high-VIX day from the scary year riskier than this low-VIX day from the calm year?" and the answer keeps being yes — not because the VIX forecast anything, but because both variables drifted with the regime.

That is a real relationship. It just is not the one you need. You do not get to trade the difference between 2021 and 2022 — you were living through one day at a time, and the question that mattered was always "is tomorrow riskier than average, given what the VIX is telling me today?" Split the data by year and that is exactly what you are measuring, because every comparison now happens between days in the same regime. The answer, four years out of seven, is worse than a coin flip.

This is the same family of mistake as confusing co-movement with prediction, one level up: the VIX and crypto risk genuinely move together across regimes, and an aggregate statistic will happily convert that co-movement into what looks like forecasting skill.

Two different traps, one lesson

It is worth putting this next to the Fear & Greed result, because the failures are not the same shape.

One lesson covers both: a single headline number is not a track record. Ask when the skill was there, not just whether it averages out positive. That is why every skill figure we publish carries a per-year breakdown beside it — not as a nicety, but because we have watched the pooled figure and the yearly figures disagree in both directions.

What this does not prove

Three questions to ask any indicator

You do not need our data to apply this. Whenever someone shows you a gauge, an index or a signal, these three questions separate a measured claim from a vibe, and almost nothing survives all three.

A useful habit: when a number surprises you in a good way, assume it is a pooling artefact or a lookahead until you have ruled both out. That instinct costs you a few hours and saves you from building on a result that was never there. It is also, uncomfortably often, the difference between a backtest and a live drawdown — ranking skill that only exists in aggregate does not show up in your account.

Why we keep publishing these

Because the alternative is asking you to take our word for it. Anyone can put a gauge on a dashboard; the question that separates a tool from a decoration is whether its numbers have ever been scored, and scored in a way that could have embarrassed the people doing the scoring.

In w4rn every series shows its measured skill, per year, next to the reading — including the ones that come out looking like this. Our own forecasts are cross-validated and Platt-calibrated, so a stated 30% is meant to be a real 30%: when we say it, those events happen about 30% of the time. If that sounds like an odd thing to aim for, that is what calibration means, and it is a much harder bar than "the backtest looked good".

Method note: figures from our univariate forward-AUC screen, combinatorial purged cross-validation, snapshot dated 8 July 2026, 1,483 observations for the VIX. Offline evidence, not a live trading claim.