What Is a Distortion Risk Measure?
Distortion risk measure is a risk statistic built by bending the probability curve before averaging. A weighting function is applied to the odds of exceeding each loss level, and value at risk and expected shortfall are both this construction with different weights.
Most risk statistics are an average of something. A distortion measure changes the probabilities before taking that average, and the curve used to change them is where all the disagreement lives.
Defined exactly
Start with the chance of exceeding each level. For every possible loss size, there is a probability the outcome is worse than that. Plotted across all levels, those probabilities form a curve running from one down to zero.
Apply a function to that curve. The function must start at zero, end at one, and never decrease. Anything else about its shape is free.
Then integrate what comes out. The result is one number, and it is a weighted average of the distribution’s quantiles where the weights were set by how steeply the function bends.
Leave the function as a straight line and you get the plain expected value. Every familiar measure in this family is a departure from that line.
Value at risk and expected shortfall are the same construction
Value at risk uses a step. The weighting function jumps from zero to one at the chosen confidence level and is flat either side, which places the entire weight on a single point of the distribution and none anywhere else.
Expected shortfall uses a ramp. The function rises linearly until it reaches the confidence level and is flat beyond, which spreads weight evenly across every outcome in the tail rather than concentrating it on the boundary.
Everything said about the two follows from those shapes. Value at risk “ignores the shape of the tail” because a step function evaluates one point. That is not a flaw someone discovered — it is arithmetic, and it was chosen.
A worked example
Take this site’s measured bar ranges. The median is 0.493 and the ninetieth percentile is 1.101, and the largest single bar recorded is 2.338.
A value-at-risk weighting at the ninety percent level reports 1.101 and stops. That is the boundary, and the measure has no opinion about anything past it.
An expected-shortfall weighting averages the whole tail beyond it. That average cannot be below 1.101, because every observation in the tail is at least that large, and it cannot exceed 2.338, because that is the largest one there is.
So the same data, under two weighting functions, produces figures that can differ by more than a factor of two. Nothing about the underlying distribution changed. Only the curve did.
And the gap widens as the tail gets fatter. The largest bar here is about 4.7 times the median, which is exactly the condition under which reporting the boundary instead of the average understates the most.
Why concavity decides coherence
A coherent risk measure has to be subadditive. Combining two positions must never score worse than holding them separately, because otherwise the arithmetic says diversification increased risk.
For this family there is a clean test. A distortion measure is coherent if and only if its weighting function is concave — bending one way throughout, never kinking back.
Expected shortfall’s ramp is concave, so it passes. Value at risk’s step is not concave anywhere, so it fails, and the celebrated counterexamples where two positions combine into a larger figure than their sum are a direct consequence of that shape.
Which reframes the whole argument. The choice between them is not a dispute about which is more accurate. It is a decision about whether the weighting function is allowed to be a step, and the subadditivity failure comes attached to that decision whether or not anyone notices.
What the weighting function is actually saying
It states how much extra attention bad outcomes deserve. A steep bend near the tail says the worst cases matter far more than their raw probability suggests.
That is a preference, not a measurement. Nothing in the data selects the curve; the institution choosing it is stating what it is willing to tolerate.
Which is why these measures came out of insurance pricing. The same construction sets premiums, where the whole job is charging more than the expected loss in proportion to how unpleasant the tail is.
And it is why two firms can report honest, correct, wildly different risk numbers on identical positions. They bent the curve differently, and both disclosed it.
The original data
On this site’s shared series: 54% of 566 ten-bar windows finished higher, the deepest drawdown is 3.76%, and the longest stretch below a prior peak runs 73 bars.
The worst ten percent of those windows is about 57 events. A value-at-risk reading describes the single window at that boundary; an expected-shortfall reading describes the average of all 57, and only the second one changes when the worst few get worse.
And 95% of bars sit below a prior peak. A tail statistic estimated from a sample where drawdown is the normal state rather than the exception needs enough observations in the tail to mean anything, which is the practical limit on every measure in this family.
Where it came from and who uses it
The idea is from decision theory. Rather than distorting outcomes through a utility curve, the dual approach distorts the probabilities and leaves the outcomes alone, which turns out to generate this entire family.
Insurance adopted it first, because a premium is exactly a reweighted expected loss and the transformations used to price catastrophe cover are distortion functions by another name.
Banking regulation moved toward it later, and the shift in market-risk rules from a value-at-risk standard to an expected-shortfall one is precisely a change of weighting function.
Trading uses it mostly without naming it. Anyone sizing against a worst-case average rather than a confidence boundary has already chosen the concave curve, whether or not the word distortion appeared anywhere.
When it fails
The characteristic failure is reporting the boundary as if it were the magnitude. A value-at-risk figure is handed to a committee, and every person in the room hears “this is what we could lose” when the statement is “we exceed this on one day in ten, and the measure is silent about by how much.” Two portfolios with identical readings can hold completely different tails, because a step weighting evaluates one point of the distribution and discards the rest by construction. The number is not wrong. It answers a narrower question than the one being asked of it, and nothing in the figure itself signals that.
A second failure is setting limits on a measure that is not subadditive, which lets exposure be split across books until the reported total falls.
A third is choosing the confidence level after seeing the answer, which turns the weighting function into a negotiating position.
A fourth is assuming coherence buys accuracy. A concave weighting is well-behaved arithmetic applied to whatever tail the sample happened to contain, and a thin tail produces a confident wrong number.
And a fifth is comparing figures across firms without checking the weighting function and the horizon behind each one.
Related
Tail risk covers the part of the distribution every weighting choice is arguing about. Deviation risk measure covers the other formal family, which measures spread instead. And downside risk covers the plainer version of the same question.
Once I saw that value at risk and expected shortfall are the same formula with a different weighting curve, every argument about which one to use stopped being a debate about statistics and started being a question about what shape of tail you care about. That is a much easier question.
— Michael Whitman
This page is educational, not financial advice. Test every idea on your own charts before risking money.