VaR backtesting compares each day's predicted 99% VaR with the next day's P&L and counts exceptions. A good model has about 1 exception per 100 days, scattered randomly. Regulators judge the count; validators also judge the pattern. This guide covers both, with the numbers a market-risk quant is expected to know.
1. What counts as an exception
An exception (or breach) is a day when the loss is larger than the VaR forecast made the day before. The harder question is which P&L you compare to. FRTB counts exceptions on two series:
- Actual P&L: what the desk really booked, including intraday trading, fees and new trades.
- Hypothetical P&L: the P&L if positions were frozen at the prior close and only the market moved. This isolates model quality from trading activity.
Using only actual P&L can hide a bad model behind lucky intraday trading; using only hypothetical P&L can hide data and booking problems. Banks track both.
2. The Basel traffic light (250 days, 99% VaR)
| Zone | Exceptions | Multiplier | Chance for a perfect model |
|---|---|---|---|
| Green | 0 to 4 | 3.00 | 89.2% |
| Yellow | 5 to 9 | 3.40, 3.50, 3.65, 3.75, 3.85 | 10.8% |
| Red | 10 or more | 4.00 | 0.03% |
The last column is worth remembering. Even a perfectly calibrated model lands in yellow about one year in nine, so yellow is a warning, not proof of failure. Red is different: 10 breaches would happen to a correct model about once in 3,000 years, so it is treated as evidence the model is wrong. Expect 2.5 breaches a year on average.
3. FRTB desk and bank-level tests
- Bank level: the internal-model multiplier starts at 1.5 and rises with 99% exceptions over 250 days: plus factors of 0.20, 0.26, 0.33, 0.38 and 0.42 for 5 to 9 exceptions, and 0.50 from 10.
- Desk level: a desk with more than 12 exceptions at 99%, or more than 30 at 97.5%, over 250 days loses internal-model eligibility and falls back to the standardised approach.
4. The Kupiec test: is the exception rate right?
The proportion-of-failures test asks whether the observed rate x/n is consistent with the target p = 1%. Under the null, the likelihood-ratio statistic is chi-square with one degree of freedom:
$$LR_{POF} = -2\ln\!\Big[(1-p)^{n-x}p^{x}\Big] + 2\ln\!\Big[\big(1-\tfrac{x}{n}\big)^{n-x}\big(\tfrac{x}{n}\big)^{x}\Big] \sim \chi^2_1$$
Worked example from the Desk2Quant lab: over 4,919 days you expect 49.2 exceptions at 1%. A model with 68 exceptions gives LR = 6.49 and p = 0.011, so it is rejected at the 5% level. A model with 43 gives LR = 0.82 and p = 0.365, so it passes. A model with 75 gives LR = 11.8 and p = 0.0006, a clear fail.
5. The Christoffersen test: do exceptions cluster?
Count how often an exception follows an exception versus a quiet day. If breaches arrive in streaks, the model reacts too slowly to changes in volatility. The independence test compares the two conditional breach probabilities. A model can pass Kupiec with exactly the right count and still fail independence because every miss falls in the same crisis month. Regulators care about this because clustered breaches happen exactly when capital matters most.
6. What 4,919 days of real data showed
The lab walks four VaR models forward from 2007 to 2026 on a rates and FX book built from real Treasury and FX history.
| Model | Exceptions (49 expected) | Kupiec p | Independence p | Verdict |
|---|---|---|---|---|
| Historical simulation (500d) | 68 | 0.011 | 0.0003 | Fails both |
| Filtered HS (500d) | 43 | 0.365 | 0.394 | Passes both |
| Parametric | 75 | 0.0006 | n/a | Fails coverage |
Plain historical simulation weighs old calm days equally with today, so after volatility jumps it keeps under-predicting and breaches in streaks. Filtered historical simulation rescales historical returns by current volatility, which fixes the clustering. That is why many desks run FHS or volatility-scaled models for limits.
7. A workflow for when you breach
- Check the data: bad price, stale curve, wrong FX rate or booking error?
- Check the P&L: does hypothetical P&L match a clean revaluation?
- Classify: data problem, model failure, or a genuine move beyond the confidence level.
- Quantify: a VaR explain showing which risk factors drove the gap.
- Escalate and document: validators expect a reasoned classification for every exception, not just the count.
8. Common mistakes
- Backtesting against P&L that includes intraday trading and fees without saying so.
- Treating yellow as a failure, or green as a pass of the whole model. Green only means the count is fine.
- Ignoring clustering and running only a count test.
- Using too short a window: with 250 days the test has little power, so validators also look at longer samples.
9. Interview answers
A model has 6 exceptions in 250 days. What do you do? "Yellow zone. A perfect model does this about 1 year in 12, so I would first rule out data and P&L issues, classify each breach, and look at clustering before touching the model. The multiplier rises by 0.50 under the Basel table."
Why run Christoffersen as well as Kupiec? "Kupiec checks the count; Christoffersen checks that breaches are independent. A model can have the right count and still cluster in a crisis."