lucid.rodeo · incerto · var-wars

The VaR file

In 1994 a bank compressed the risk of a whole trading book into one dollar figure and gave the method away for free. People have been arguing ever since about whether that number was a service or a sedative. This is the file: the record, both defenses at full strength, and what the regulators actually did.

Contested Thirty years of argument about one number — the record, both defenses, and what the regulators actually did; no verdict sold. Every claim below carries its own chip: Factdocumented in the cited record  ·  Contesteda live dispute, both sides given  ·  Claimasserted, not settled here. the fence →
RiskMetrics launch Oct 1994
technical document, 4th ed. Dec 1996
the exchange Apr 1997
the academic warning May 2001
the stress test 2008
the retrospective Jan 2009
FRTB adopted Jan 2016 (rev. 2019)
the move 99% VaR → 97.5% ES
the file stays open — nothing below declares a winner

1 · The number

FactValue at Risk is a quantile wearing a dollar sign. When a desk reports a one-day 99% VaR of some figure, it is saying: on ninety-nine days in a hundred, by our model, the loss stays under this line. The number marks where the tail begins. About how bad the hundredth day gets — the size of what lies beyond the line — it is silent by construction, not by accident. Both sides of the war below agree on this definition; the war is about what the silence costs.

FactThe number's home is J.P. Morgan's RiskMetrics, launched in October 1994 — and the launch is a large part of why VaR conquered the industry: Morgan gave the method away free over the Internet, technical document and daily-updated data together, so the framework became everyone's default rather than one bank's edge. The definitive statement is the Fourth Edition of the Technical Document (Morgan/Reuters, December 1996). Worth filing from that same document: Morgan's own warning label, which said in plain terms that no amount of analytics replaces experience and judgment, and that the tool guarantees nothing. The parents of the number posted the caveat themselves, inside the definitive edition.

FactThe regulators folded it in. The Basel Committee went as far as letting banks set their market-risk capital from their own internal VaR models — the arrangement Nocera's retrospective walks through, and the one the 2016 standard (below) eventually replaced. For roughly two decades, the contested number was not just a dashboard reading; it was the yardstick behind bank capital.

FactThe rival measure, for later: Expected Shortfall. Where VaR reads off the loss at the quantile, ES averages every loss at and beyond it — the mean of the worst 100p% of outcomes — so it sees the size of the tail that VaR ignores by design. Acerbi & Tasche's 2002 paper made the definitional case and added the technical one: ES is coherent — it satisfies subadditivity, so a diversified book can never measure as riskier than the sum of its parts, a property VaR can violate.

2 · The 1997 exchange

FactThe exchange happened, in print, in Derivatives Strategy, April 1997: Philippe Jorion's piece defending VaR, and Nassim Taleb's reply against it (hosted on his site under a longer title; Jorion's side survives in an archived capture of his university page, which has since gone dark). The trigger was Jorion's criticism of a Taleb interview from the magazine's December/January issue. Everything below is paraphrase — both pieces are short, and worth reading whole.

ContestedTaleb's charge, at full strength. Three planks. First, tail blindness: the whole apparatus stands on measuring the probability of rare events — precisely the region where probabilistic measurement is weakest — and a quantile says nothing about the size of what lies past it, which is where ruin lives; he argued the approach had already been falsified in practice by a string of blow-up years he lists from 1985 to 1995. Second, false confidence is not a side effect but the product: a scientific-looking number invites people without tail experience to lean on it and take risks they are not equipped to judge — the instrument manufactures the very confidence that does the damage. Third, he pushed the charge all the way to charlatanism — his borrowed definition: concealing a poor understanding behind mathematical smoke — and said he would rather suspend the method as potentially dangerous malpractice than refine it.

ContestedJorion's defense, at full strength. VaR never claimed to describe the worst case — its stated job is an estimate of the range of possible gains and losses, and blaming a 99% quantile for the contents of the other 1% is blaming the ruler for the mountain. Its real product is discipline: a structured process for thinking critically about risk, in which getting to the number can matter as much as the number itself — his exhibit was Orange County, his standing example of the number as a communication and discipline device. His image for the tool: a wobbly speedometer — imprecise, yes, and still better than driving with no gauge at all, because the alternative actually on offer was not a better number but no number. On the strongest insult, his reply was simply that calling it charlatanism was premature.

FactThe concession that keeps the file honest. The exchange was not a clean collision: Jorion himself conceded the gaming point — traders can construct positions that look quiet on VaR while quietly loading the tail (short-option structures of the kind Leeson ran). The defense granted one of the attack's genuine observations and stood anyway. That is why this file can be kept without a verdict: the two sides are not even disagreeing about all of the facts, only about what follows from them.

3 · The stress test of 2008

FactThe warning predates the crisis — that is the load-bearing fact. In May 2001, in the formal comment round on Basel II, seven academics — Daníelsson, Embrechts, Goodhart, Keating, Muennich, Renault and Shin — filed an academic response (LSE Financial Markets Group Special Paper No. 130) arguing, in paraphrase: risk is endogenous — when every institution holds the same VaR yardstick, the measure stops observing the system and starts driving it, and can induce crashes that would not otherwise occur; the statistical forecasts in use were biased and understated joint downside risk; and capital regulation is inherently procyclical, and rules built on this yardstick would exacerbate it significantly — harmonizing behavior in a crisis and amplifying exactly what they were meant to damp. It was in the Basel Committee's hands, as a formal comment, seven years before 2008.

Read one way, it is a critique of VaR-based regulation — one yardstick imposed on everyone — even more than of the statistic itself. That reading is ours, not the authors'; the paper is a click away.

FactThe documented retrospective. January 2009, rubble still warm: Joe Nocera's Risk Mismanagement (The New York Times Magazine, published January 2, 2009; print January 4 — an archived capture of the article) put the question the way this page keeps it: did the measures themselves fail, or did the failure lie in ignoring what they said? The piece traces the number to the JPMorgan quants, locates its charm in compressing risk to a single dollar figure, and records that the one risk sitting outside VaR's field of view was the largest of all — a full financial meltdown. And it is genuinely two-sided: Einhorn's image of a safety device that works right up until the crash sits alongside the pro-VaR exhibit of Goldman reacting to VaR-triggered warnings in late 2006, with Taleb present and still prosecuting.

Perhaps the best single document of the argument mid-explosion — and it declines to convict either side. That ranking is ours; the capture is linked above.

4 · What the regulators did

FactIn January 2016 the Basel Committee's Fundamental Review of the Trading Book — Minimum capital requirements for market risk, BCBS d352; revised January 2019 as d457, keeping the same design — moved market-risk capital from VaR at 99% to expected shortfall at the 97.5th percentile. The Committee's own stated rationale, paraphrased from the standard: shifting to ES gives a more prudent capture of tail risk and better capital adequacy in periods of significant market stress. Backtesting still runs against desk-level VaR at both the 97.5th and 99th percentiles — the old number keeps a job as the audit gauge.

Read the shift carefully: it is a documented regulatory move toward a tail-aware measure — the quantile that ignores the tail's size was replaced, for capital purposes, by the average that reads it. It is not a declared winner of the 1997 exchange. The BCBS adopted a tail-sensitive statistic, not Taleb's rhetoric — and a benchmark-and-discipline reading of the new number is as available to Jorion's side as it ever was for the old one.

5 · Still contested

ContestedDid VaR fail, or was it misused? One telling: the number did its stated job — a quantile of ordinary days — and the failure was institutional, the treating of a fair-weather gauge as a worst case; that is the misuse reading, and it stands close to Jorion's ground and to half of Nocera's frame. The other telling: a measure whose stated job excludes the events that matter structurally invites that misuse — the false confidence is not an accident of careless users but the instrument working as designed; that is Taleb's ground and the frame's other half. The record above is consistent with either. This page files both and buys neither.

ContestedIs expected shortfall estimable where it matters most? ES is a tail average, and averages need moments. Estimating the mean of the worst losses requires the tail to be tame enough that this mean exists and settles at practical sample sizes — and fat-tailed regimes are precisely where that assumption buckles. Elsewhere on this wing, the lying average shows a sample mean refusing to settle across a million draws, and the tails page shows the tail exponent deciding which moments exist at all. If the exponent sits low enough, the quantity ES is hired to estimate converges cruelly slowly, or does not exist — the estimator inherits the disease it was brought in to diagnose. Whether real trading books live in the tame zone is an open technical question, argued in live journals — one side holds that at the FRTB's 97.5% line the tail is tame enough for the average to settle, the other that the exponent sits too low; the honest chip here is the gold one.

ClaimThat the argument moved the regulators. It is tempting to read the 2016 rule as the 1997 exchange winning slowly. Nothing in this file documents that causation: the standard states its own reasons — prudence about tail risk under stress — and cites neither combatant. The line from the debate to the rule is an assertion, so it wears the dashed chip, and can be discounted accordingly.

Sources — the whole file, in order. J.P. Morgan (later Morgan/Reuters), RiskMetrics — Technical Document, first released with the October 1994 launch, definitive Fourth Edition December 17, 1996. The Jorion–Taleb debate, Derivatives Strategy, April 1997: Philippe Jorion, In Defense of VAR (archived capture); Nassim Taleb, Against VAR. Jón Daníelsson, Paul Embrechts, Charles Goodhart, Con Keating, Felix Muennich, Olivier Renault and Hyun Song Shin, An Academic Response to Basel II, LSE Financial Markets Group Special Paper No. 130, May 2001. Joe Nocera, Risk Mismanagement, The New York Times Magazine, published January 2, 2009 (print January 4) — archived capture. Basel Committee on Banking Supervision, Minimum capital requirements for market risk, BIS, 14 January 2016 (BCBS d352; revised January 2019, BCBS d457). Carlo Acerbi and Dirk Tasche, Expected Shortfall: a natural coherent alternative to Value at Risk, Economic Notes 31(2), 2002. Living authors throughout — every position above is paraphrase, attributed; none of it is their prose.

the wing →