This report is built by sorting twenty decisions on their outcome and keeping the top three. That is hindsight, applied deliberately, and it is the one thing the decision record forbids anywhere inside a reason. It is legitimate here only because the report is labelled as what it is: an account of what happened after the fact, not an account of why anything was decided.
Three winners out of twenty is not evidence of skill and this note is not offered as any. The full scorecard is 10 helped, 8 hurt, 2 too small to tell, a net of 22bps a year against 2.6bps of turnover cost, and an information ratio of 0.27 whose standard error on sixty monthly observations is roughly 0.45. Selecting the best three from any twenty coin flips produces three impressive coin flips. The question worth asking of each entry below is not whether it made money but whether the reasoning tabled on the day would have looked sound to a trustee who did not yet know the answer.
| Meeting | Earned | The line that earned it | Its rank among the signal readings |
|---|---|---|---|
| bps | |||
| 31 March 2022 | 106 | US equity, 69bps | 5 of 9 |
| 31 December 2025 | 100 | Commodities, 57bps | 8 of 9 |
| 31 December 2021 | 79 | Commodities, 27bps | 3 of 9 |
In two of the three, the line that produced most of the money was not a line the signal was loud about, and in one of those it was a position the meeting did not trade at all. That is not a criticism of the decisions. It is the same finding the systematic evidence reports from the other direction: with a composite out-of-sample R² of (1.61)% against the historical mean, there is no reason to expect the loudest signal to be the one that pays, and the record shows it was not.
| Input | Reading |
|---|---|
| Regime, on the vintages available then | overheat — growth expansion, inflation high, policy accommodative |
| Three loudest composite signals | US investment grade -0.83, Commodities +0.82, Developed ex-US +0.51 |
| The two desks | disagreed, 641bps apart at the widest line |
| Ex-ante tracking error | 93bps before truncation, 93bps after, against a 200bps budget |
| Binding constraint | drawdown_stress |
| Compliance on the allocation adopted | PASS |
| Turnover and its cost | 10.46pp one-way, costing 0.80bps |
The regime read on the vintages available at this date was expansion growth with high inflation and accommodative policy, classified overheat. The composite signal was strongest on US investment grade at -0.83, Commodities at +0.82, Developed ex-US at +0.51. The reconciled allocation moved T-bills 7.2pp; US Treasury duration (5.7)pp; US equity (2.8)pp. The binding constraint at this meeting was drawdown_stress.
That paragraph is the reason as recorded at the meeting. It was built only from the readings above, it names no date later than the meeting, and it contains no reference to the outcome. That is asserted mechanically across all twenty entries by tests/check_hindsight.py.
| Line | Active weight held | Line return over the quarter | Contribution |
|---|---|---|---|
| bps | % | bps | |
| US equity | (428) | (16.11) | 69.0 |
| Developed ex-US | 316 | (13.22) | (41.7) |
| US Treasury duration | (688) | (4.47) | 30.7 |
| US investment grade | (304) | (8.40) | 25.6 |
| Emerging markets | (152) | (10.43) | 15.9 |
| Listed real estate | 62 | (15.39) | (9.6) |
| Sum of the six largest | 89.9 |
A first-order attribution: the active weight adopted at the meeting, held constant, times each line's compounded return over the quarter. The realised active return differs because weights drift with returns inside the quarter and because the arithmetic compounds. The gap between the two is the interaction term and is not attributable to any single line.
The largest contributor was the US equity underweight, worth 69bps as the line fell (16.11)% over the quarter. The composite on US equity that day was -0.27, which was the fifth loudest of nine readings. The two signals the desk cited as strongest, US investment grade at (0.83) and commodities at +0.82, contributed 26bps and less respectively.
The decision was also wrong on a line. The developed ex-US overweight, which the composite ranked third at +0.51, cost (42)bps. The meeting made money net because the de-risking was broad and the quarter was broadly negative, not because the ranking was right.
What this decision does show is the constraint hierarchy working as designed. The combined desk view implied 93bps of tracking error and the binding constraint was the ex-ante drawdown stress test, not the tracking-error budget. The office moved 10.5pp of the portfolio, the largest single quarter of turnover in the record, for 0.80bps of cost. That is the implementation desk's cost vector doing its job.
Lines that detracted at this meeting: Developed ex-US (41.7)bps, Listed real estate (9.6)bps, US high yield (5.7)bps. A decision that made money is not a decision that was right about everything, and an outcome report that shows only the winners inside a winner is the same selection error one level down.
| Input | Reading |
|---|---|
| Regime, on the vintages available then | overheat — growth expansion, inflation above_target, policy neutral |
| Three loudest composite signals | Emerging markets +0.71, US equity -0.67, Developed ex-US +0.57 |
| The two desks | disagreed, 1300bps apart at the widest line |
| Ex-ante tracking error | 78bps before truncation, 81bps after, against a 200bps budget |
| Binding constraint | drawdown_realised |
| Compliance on the allocation adopted | PASS |
| Turnover and its cost | 5.41pp one-way, costing 0.45bps |
The regime read on the vintages available at this date was expansion growth with above_target inflation and neutral policy, classified overheat. The composite signal was strongest on Emerging markets at +0.71, US equity at -0.67, Developed ex-US at +0.57. The reconciled allocation moved Developed ex-US 4.3pp; US equity (2.2)pp; Emerging markets (1.8)pp. The binding constraint at this meeting was drawdown_realised.
That paragraph is the reason as recorded at the meeting. It was built only from the readings above, it names no date later than the meeting, and it contains no reference to the outcome. That is asserted mechanically across all twenty entries by tests/check_hindsight.py.
| Line | Active weight held | Line return over the quarter | Contribution |
|---|---|---|---|
| bps | % | bps | |
| Commodities | 193 | 29.47 | 57.0 |
| US equity | (576) | (4.37) | 25.2 |
| Emerging markets | 138 | 3.80 | 5.3 |
| Developed ex-US | 375 | 1.15 | 4.3 |
| T-bills | 197 | 0.85 | 1.7 |
| US investment grade | (303) | (0.38) | 1.2 |
| Sum of the six largest | 94.5 |
A first-order attribution: the active weight adopted at the meeting, held constant, times each line's compounded return over the quarter. The realised active return differs because weights drift with returns inside the quarter and because the arithmetic compounds. The gap between the two is the interaction term and is not attributable to any single line.
This is the clearest case in the record of a good outcome that the decision did not cause. The largest contributor by a distance was commodities, worth 57bps as the line returned 29.47% over the quarter. The meeting did not trade commodities. The position was carried in from earlier decisions and the composite on that line was not among the three the desk cited.
What the meeting actually decided was to add 4.3pp to developed ex-US, cut US equity 2.2pp and cut emerging markets 1.8pp. The US equity underweight earned 25bps and was well supported by the second-loudest signal. The developed ex-US addition earned 4bps. And the emerging markets cut looks at first like it ran against the loudest signal on the table, emerging markets at +0.71. It did not. The line had drifted to 15.2% on its own returns, above the Quantitative desk's target of 14.4% and above the Macro desk's 12.5%, so both desks implied a trim and the sale was drift correction rather than a reversal of the view. This is worth stating because a column headed "what moved" invites exactly the wrong reading: a sale can be a signal-consistent overweight being pulled back to target.
A trustee is entitled to ask why the office sold the line its own model liked most. The answer is the pre-committed reconciliation rule, applied without exception across all twenty meetings, and the office would rather report a rule producing an awkward result than a rule that was quietly suspended when it did.
| Input | Reading |
|---|---|
| Regime, on the vintages available then | overheat — growth expansion, inflation high, policy accommodative |
| Three loudest composite signals | US investment grade -0.56, Listed real estate +0.49, Commodities +0.49 |
| The two desks | disagreed, 800bps apart at the widest line |
| Ex-ante tracking error | 87bps before truncation, 28bps after, against a 200bps budget |
| Binding constraint | drawdown_stress |
| Compliance on the allocation adopted | PASS |
| Turnover and its cost | 1.21pp one-way, costing 0.07bps |
The regime read on the vintages available at this date was expansion growth with high inflation and accommodative policy, classified overheat. The composite signal was strongest on US investment grade at -0.56, Listed real estate at +0.49, Commodities at +0.49. The reconciled allocation moved Developed ex-US 1.2pp; US equity (1.2)pp. The binding constraint at this meeting was drawdown_stress.
That paragraph is the reason as recorded at the meeting. It was built only from the readings above, it names no date later than the meeting, and it contains no reference to the outcome. That is asserted mechanically across all twenty entries by tests/check_hindsight.py.
| Line | Active weight held | Line return over the quarter | Contribution |
|---|---|---|---|
| bps | % | bps | |
| Commodities | 108 | 25.41 | 27.5 |
| Emerging markets | (146) | (7.57) | 11.0 |
| Developed ex-US | 123 | (6.46) | (7.9) |
| US investment grade | (84) | (8.38) | 7.0 |
| US Treasury duration | (107) | (6.37) | 6.8 |
| US equity | (94) | (4.61) | 4.3 |
| Sum of the six largest | 48.7 |
A first-order attribution: the active weight adopted at the meeting, held constant, times each line's compounded return over the quarter. The realised active return differs because weights drift with returns inside the quarter and because the arithmetic compounds. The gap between the two is the interaction term and is not attributable to any single line.
The largest contributor was Commodities, worth 27bps, and it ranked 3 of nine among the signal readings that day. The position was small: the whole meeting moved 1.2pp and cost 0.07bps.
The more interesting fact about this meeting is that it was principally a de-risking. Ex-ante tracking error fell from 87bps to 28bps, the largest single reduction in the record, because the ex-ante drawdown stress test bound and truncated the combined view. The office earned 79bps in a quarter when the benchmark fell -5.01% largely by holding less of everything that was falling.
The developed ex-US overweight, ranked fourth on the composite at +0.46, detracted (8)bps. As in the other two entries, the ranking was partly right and partly wrong and the net was positive.
Lines that detracted at this meeting: Developed ex-US (7.9)bps, Listed real estate (3.9)bps, US high yield (2.6)bps. A decision that made money is not a decision that was right about everything, and an outcome report that shows only the winners inside a winner is the same selection error one level down.
These three decisions produced 285bps between them. The other seventeen produced -177bps. The programme as a whole returned 22bps a year before the cost of running the office, on an information ratio that cannot be separated from zero at this sample size.
A record in which three decisions carry the entire result, and in which the line that earned the money was usually not the line the signal was loudest about, is a record of a programme that got lucky in a few quarters rather than one that found something. The office does not have twenty independent observations here; it has one tightening cycle and one recovery. The recommendation for the coming year is policy weights, and nothing in this note argues otherwise.
The one decision-useful finding is the mirror image of this note. Of the eight decisions that hurt, all eight carried a Treasury duration position and the sign flipped halfway through the record. That is one error made twice on the line with the most negative out-of-sample R², and it is fixable in a way that three good quarters are not repeatable.