There is a particular kind of model failure that only reveals itself in a crisis: the model that is right 99% of the time and catastrophically silent about the other 1%. The 2007–09 financial crisis exposed value-at-risk as exactly that kind of model. Banks held capital against a quantile of the loss distribution while the losses that mattered lived beyond it, in a region VaR is constructed to ignore. The regulatory response took years to crystallise, but it was unambiguous: the Basel Committee's Fundamental Review of the Trading Book, finalised in its revised form in January 2019, replaced 99% VaR with a 97.5% expected shortfall, calibrated to a period of significant financial stress, as the core internal-model measure for market-risk capital (BCBS, 2019a; BCBS, 2019b).
Implementation timelines have shifted repeatedly across jurisdictions — the EU consulted during 2025 on further delay amid uncertainty about the US timetable, with proposals leaving parts of the internal-models framework to later dates than the standardised elements (KPMG, 2025). But the methodological verdict is settled, and it applies well beyond banks. Any institution still managing market risk primarily through an unstressed historical VaR — including insurers and asset managers with no FRTB obligation at all — is running its risk framework on a measure the last two decades have specifically discredited for tail purposes.
The business problem: three failures of the historical-VaR habit
The traditional setup — a one-sided quantile of P&L, estimated from a rolling window of recent history — fails in three distinct ways, and it is worth separating them because they have different remedies.
It ignores tail severity. VaR at level α answers: what loss is exceeded with probability 1−α? It is indifferent between a distribution where exceedances are modest and one where they are ruinous. Expected shortfall repairs exactly this.
It is calibrated to calm. A rolling recent window mechanically produces low risk estimates after quiet years — the moment of maximum complacency — and spikes after the crash has already happened. Stressed calibration repairs this.
It assumes yesterday's dependence. Historical windows dominated by normal markets embed normal-market correlations. In stress, diversification that existed on paper evaporates as assets fall together. Explicit tail-dependence modelling repairs this — and no change of risk measure alone does.

The measures, precisely and in plain language
For a loss random variable L, value-at-risk at confidence level α is the quantile:
VaR_α = inf { ℓ : P(L > ℓ) ≤ 1 − α }
— the smallest loss threshold that losses exceed with probability at most 1−α. Expected shortfall at the same level is (for continuous distributions):
ES_α = E[ L | L ≥ VaR_α ]
— the average loss on the days the threshold is breached. In plain terms: VaR locates the doorway to the tail; ES walks through it and reports the mean of everything inside.
Two properties explain the regulatory choice. First, ES is subadditive: the ES of a combined portfolio never exceeds the sum of the parts' ES, so the measure never punishes diversification — a coherence property VaR famously lacks for non-elliptical distributions (McNeil, Frey & Embrechts). Second, the calibration was chosen for continuity: for a normal distribution, 97.5% ES (≈ μ + 2.34σ) almost coincides with 99% VaR (≈ μ + 2.33σ), so the Committee's move preserved capital levels for thin-tailed books while automatically demanding more wherever tails are fat (BCBS, 2019b) — which is to say, wherever it matters. The FRTB pairs the measure with calibration to a 250-day stressed period and with liquidity horizons of 10 to 120 days by risk-factor class, replacing the fiction that every position exits in ten days (BCBS, 2019a).
One honest caveat belongs in every ES discussion: backtesting. ES is statistically harder to validate directly than VaR — which is why the FRTB itself backtests the model's VaR at the 97.5% and 99% levels through exception counting, using the standard battery of unconditional and conditional coverage tests, while capitalising on ES (BCBS, 2019b). A pragmatic validation design does the same: exception tests on quantiles, complemented by benchmarking and sensitivity analysis on the tail average.
Tail dependence: the correlation that shows up only in crises
The subtler failure is multivariate. Between two risks with uniform marginals U₁, U₂, the coefficient of upper tail dependence is:
λ_U = lim(q→1) P( U₂ > q | U₁ > q )
— the probability that one risk has an extreme outcome given the other does, in the limit of ever more extreme thresholds. The Gaussian dependence structure that underlies most correlation-matrix thinking has λ_U = 0 for any correlation below one: extremes are asymptotically independent, and joint crashes are essentially impossible in the model even when they are routine in markets. A Student-t copula, by contrast, produces strictly positive tail dependence governed by its degrees-of-freedom parameter — a documented, testable way to encode the empirical fact that diversification fails precisely when it is needed (McNeil, Frey & Embrechts). Regular readers will recognise the same machinery deployed on climate perils elsewhere in this series; the mathematics of "everything goes wrong together" is peril-agnostic.

A practical workflow: upgrading a legacy VaR framework
For a multi-asset investment book still on rolling historical VaR, a proportionate upgrade — of the kind Quantica Risk's market-risk framework development follows — runs:
- Compute ES alongside VaR on the existing engine at 97.5%; the incremental cost is trivial, and the ES/VaR ratio per desk is itself a diagnostic — ratios well above the normal-distribution benchmark flag fat-tailed books.
- Select and document a stressed window by searching history for the 250-day period maximising the portfolio's modelled loss with current weights — and record the rationale, because window choice is a governable assumption, not a technicality.
- Test the dependence structure: fit Gaussian and Student-t copulas to the key risk-factor pairs; if the t-copula's degrees of freedom come out low, joint-tail risk is materially understated by the correlation matrix.
- Report three numbers per desk — unstressed ES, stressed ES, and the joint-tail sensitivity — instead of one VaR; the spread between them is the risk conversation the old framework suppressed.
- Backtest what is backtestable: exception counts on VaR quantiles, stability monitoring on ES.
Implications for risk leaders
First, if your institution reports a single unstressed VaR to its risk committee, the committee is seeing the most flattering available description of the book; mandate ES and a stressed variant alongside it regardless of regulatory scope. Second, interrogate diversification claims: ask explicitly what dependence assumption produced the netting benefit, because a Gaussian assumption manufactures diversification in the tail by construction. Third, treat the stressed-window choice and the copula family as named, owned model assumptions with sensitivity analysis attached — that is where the real model risk in a modern market-risk framework now sits, not in the arithmetic of the quantile.

Conclusion
The move from VaR to expected shortfall is often described as a technical refinement. It is better understood as a philosophical correction: risk management exists for the tail, so the headline measure should describe the tail rather than merely locate it. Stressed calibration and honest tail-dependence modelling complete the correction — the first anchors the measure to conditions that matter, the second stops the framework inventing diversification that stress will revoke. Later articles in this pillar extend the same logic to liquidity risk under stress, and to regime-switching models in which correlations, volatility and liquidity shift together — because the deepest lesson of the last two decades is that the state of the world is itself a risk factor.
About the author. Jonas Osman is a risk-modelling and financial-risk professional and the founder of Quantica Risk, an AI-driven modelling company focused on insurance, banking, climate risk, actuarial analytics, and model validation.
Quantica Risk develops transparent, data-driven modelling frameworks for financial institutions and risk-sensitive businesses. To discuss market risk, stress testing, or model validation, visit the Quantica Risk website. [website — to be inserted when verified]
Series links
- Next in this pillar: Liquidity Risk Under Stress: When Market and Funding Liquidity Spiral Together (Article D2) — the risk dimension ES still leaves out.
- Related: Pricing the Tail: Extreme-Value Theory, GPD, and the Market Price of Catastrophe Risk (Article A3) — the same tail problem, seen from the insurance side.

References
- Basel Committee on Banking Supervision (2019a). Minimum capital requirements for market risk (revised standard, January 2019). https://www.bis.org/bcbs/publ/d457.htm
- Basel Committee on Banking Supervision (2019b). Explanatory note on the minimum capital requirements for market risk. https://www.bis.org/bcbs/publ/d457_note.pdf
- Basel Committee on Banking Supervision (2013). Fundamental review of the trading book: A revised market risk framework (consultative document). https://www.bis.org/publ/bcbs265.pdf
- Bank Policy Institute (2023). Why is the FRTB Expected Shortfall calculation designed as it is? https://bpi.com/why-is-the-frtb-expected-shortfall-calculation-designed-as-it-is/
- SIFMA (2021). The Fundamental Review of the Trading Book (FRTB): An introductory guide. https://www.sifma.org/news/blog/the-fundamental-review-of-the-trading-book-frtb-an-introductory-guide
- KPMG (2025). Fundamental review of the trading book: An overview (implementation timelines and 2025 EU consultation). https://assets.kpmg.com/content/dam/kpmgsites/in/pdf/2025/07/fundamental-review-of-the-trading-book-an-overview.pdf
- McNeil, A. J., Frey, R., & Embrechts, P. Quantitative Risk Management: Concepts, Techniques and Tools (rev. ed.). Princeton University Press. (Coherent risk measures, copulas and tail dependence.)
