The Endowment should hold policy weights for the year to 30 June 2027, take no intentional active risk, and rebalance only on a corridor breach. The office further recommends that the Board be asked two amendment questions under IPS 2.3, because two of the objectives in the Statement cannot currently be met and neither can be fixed inside the portfolio.
Three findings force this, reached by three different desks. Two of the three are genuinely independent of each other: a valuation gap and a breadth-and-cost arithmetic share no input. The third, the out-of-sample evidence, updates from the same replication literature the second does, so it should be counted as corroboration rather than as a third vote. The policy portfolio is priced to earn 6.08% over ten years against the 8.10% the spending rule requires, a shortfall of (202)bps, and assembling the single most optimistic published forecast for every line still reaches only 7.43%. The tactical programme's expected information ratio, on the fundamental law applied to this mandate's own breadth and constraints, is 0.021, worth 4.2bps a year against 6.4bps of turnover cost. And of forty signal and line combinations tested out of sample, nine beat the expanding historical mean, which is where the replication literature says the prior should have been.
The five years behind this report support the same conclusion. Twenty quarterly decisions produced 22bps of active return a year against 2.6bps of turnover cost, an information ratio of 0.27 on sixty monthly observations. The standard error on a Sharpe ratio at that sample size is ±0.45, which is wider than the entire result. 10 decisions helped, 8 hurt and 2 were too small to tell. The programme is not distinguishable from having done nothing, and doing nothing is cheaper.
The Statement anticipated this. IPS 3.3 states that the return objective and the drawdown limit are in tension by construction and that any recommendation must say which it is giving ground on. This one gives ground on the return objective. It does so because the drawdown limit is rank 3 in the constraint hierarchy and the return objective is rank 5, and because the evidence says the risk required to close the gap is not available at an acceptable drawdown.
| Question | IPS | The finding | What the office recommends |
|---|---|---|---|
| The spending rule cannot be funded from the policy portfolio | 3.2, 2.3 | The policy portfolio is priced to earn 6.08% against 8.10% required. The gap is (202)bps on the median of seven houses and (67)bps even on the most optimistic forecast available for every line simultaneously. | Reduce the spending rate, accept a lower real corpus, or amend the policy portfolio. All three are Board decisions. None is available to this office. |
| The drawdown limit is inconsistent with the policy portfolio | 3.3, 4.1, 2.3 | The policy portfolio breached the (20.00)% limit in this window, reaching (22.46)% in September 2022, and the risk desk's ex-ante estimate for policy weights is (21.60)%. The limit is breached by the Board's own policy portfolio in a normal cycle, without any tactical position. | Either widen the limit, or change the policy portfolio so that it can respect it. The office cannot deliver a (20.00)% ceiling from a 70% equity allocation. |
Both are escalated rather than absorbed. IPS 2.3: "Where analysis shows an objective in this Statement to be unattainable, the finding is escalated to the Board as an amendment question. It is not resolved inside the portfolio by taking risk the Statement does not permit."
Six desks, staffed at the outset rather than when the work stalled. A desk earned its place by owning work that could run without waiting on another desk and that something outside it could prove wrong. Each tabled a written paper and a check that returns pass or fail without a judgement call. The papers are in desks/ and the checks in tests/, and both are part of this submission rather than working material behind it.
| Desk | What it owned | Its check, which can fail | Result |
|---|---|---|---|
| Capital Markets | Long-horizon return forecasts for all nine lines, and what the policy portfolio is therefore priced to earn. | Numbers reconcile to seven published house forecasts a reader can open. | 5 of 5 assertions |
| Systematic | What predicts returns out of sample, whether volatility management works, and what the fundamental law caps this programme at. | Every figure carries a source and a VERIFIED or RECALLED status; the fundamental-law arithmetic recomputes. | 26 of 26 checks |
| Implementation & Operations | Transaction costs on the named vehicles, corridor design, and the reporting standard. | Cost assumptions reconcile to issuer-published 30-day median spreads under SEC Rule 6c-11. | 53 of 53 checks |
| Quantitative | An allocation from the signals alone, reached without sight of the macro view. | Out-of-sample statistics reported in full including every negative cell. | look-ahead and range assertions at all 20 dates |
| Macro | An allocation from the regime and from what is already priced, reached without sight of the model output. | Every deviation from consensus carries a named falsifier and a date. | 688 assertions |
| Risk | The compliance test, as code, run against every proposed allocation. | The test shown rejecting a non-compliant allocation on every binding constraint. | 13 of 13 mutants died |
The Quantitative and Macro desks are the pair that must not see each other's work. Independence here was structural rather than promised, and the trustees are entitled to know exactly what made it so. The two desks were commissioned in the same instruction and ran concurrently, so neither desk's output existed when the other began. Their briefs were disjoint: neither contained any conclusion, number or framing from the other, and both are reproduced verbatim in the evidence appendix so that this can be checked rather than taken on trust. Each was barred by name from reading the other's files, and their outputs were written to separate directories. Reconciliation happened afterwards, in this office, and both pre-reconciliation drafts are tabled unchanged.
What the two desks did share is a data layer, and that is the honest limit on the claim. Both read the same prices and the same point-in-time macro vintages through the same module. Where they agree, a shared input may be doing the work in both, and the reconciliation below says so where it applies.
Reconciliation, portfolio construction and the writing were retained by the Chief Investment Officer. Those need everything in one head, which is the reason they are not a desk.
Performance is presented against the benchmark and never in isolation, to the Global Investment Performance Standards for Asset Owners (2020 edition). The benchmark is the policy portfolio at the IPS 4.1 weights, rebalanced monthly. That rebalancing frequency is a disclosure item under provision 24.C.27 rather than an implementation detail, and it is not neutral. Over this window a never-rebalanced policy portfolio would have returned 8.56% against the 8.26% of the monthly-rebalanced blend, so the benchmark this report uses is 29.7bps a year easier to beat. The choice flatters the fund. It was made before the comparison was run and is disclosed here rather than revised, and the active return would be lower against the harder benchmark.
| Period | Fund | Benchmark | Active |
|---|---|---|---|
| % | % | bps | |
| Q1 FY26 | 5.80 | 5.97 | (17) |
| Q2 FY26 | 2.80 | 2.67 | 13 |
| Q3 FY26 | 0.88 | (0.12) | 100 |
| Q4 FY26 | 9.26 | 10.08 | (81) |
| FY2026 | 19.88 | 19.61 | 27 |
GIPS provision 24.A.1.j requires the three-year annualised ex-post standard deviation using monthly returns, for the benchmark as well as for the total fund, as of each annual period end. It is the requirement in-house reporting most often misses, and it is given here for both.
| Period | Fund | Bmk | Active | Fund σ | Bmk σ | TE | IR | Fund DD | Bmk DD |
|---|---|---|---|---|---|---|---|---|---|
| % | % | bps | % | % | bps | % | % | ||
| 1 year | 19.88 | 19.61 | 27 | 9.33 | 9.58 | 85 | +0.24 | (4.52) | (4.71) |
| 3 years | 15.56 | 15.54 | 2 | 9.40 | 9.74 | 77 | -0.02 | (7.91) | (8.37) |
| 5 years | 8.60 | 8.26 | 34 | 11.77 | 12.33 | 94 | +0.27 | (20.70) | (22.46) |
| FY2022 | (11.08) | (12.97) | 189 | 11.72 | 12.04 | 102 | +2.06 | (14.68) | (16.39) |
| FY2023 | 10.10 | 10.78 | (68) | 16.65 | 17.71 | 120 | -0.66 | (11.47) | (12.23) |
| FY2024 | 14.24 | 13.76 | 48 | 11.66 | 12.12 | 78 | +0.48 | (7.91) | (8.37) |
| FY2025 | 12.68 | 13.36 | (68) | 7.50 | 7.84 | 70 | -0.90 | (2.88) | (2.64) |
| FY2026 | 19.88 | 19.61 | 27 | 9.33 | 9.58 | 85 | +0.24 | (4.52) | (4.71) |
| Since inception | 8.60 | 8.26 | 34 | 11.77 | 12.33 | 94 | +0.27 | (20.70) | (22.46) |
Returns of less than one year are not annualised (GIPS 22.A.9). Fiscal years end 30 June. Returns are net of the transaction costs modelled in taa/costs.py and gross of the 0.40% cost load in the IPS 3.2 return requirement, which is an office and custody cost rather than a trading cost. Negatives are shown in parentheses throughout.
The one-year number and the five-year number disagree, and the three-year number disagrees with both. The fund added 27bps of active return over FY2026, 2bps a year over three years, and 34bps a year over five. None of these is distinguishable from zero. On sixty monthly observations the standard error of the information ratio is roughly 0.45, so a measured 0.27 sits comfortably inside the interval that also contains no skill at all, and comfortably inside the one that contains twice the skill. The right reading of three numbers that disagree is that the sample is too short to separate them.
This is the most consequential finding in the report and it has nothing to do with the tactical programme. The benchmark, which is the Board's own policy portfolio held passively, fell (22.46)% peak to trough into September 2022. The limit at IPS 3.3 is (20.00)%. The policy portfolio breached the Board's limit in an ordinary cycle, with no tactical position taken and nothing unusual done. The fund fell (20.70)%, which is less, and still a breach.
Every historical decision in this record is mechanical. A pre-committed rule read the point-in-time inputs available on the meeting date and produced an allocation. The trustees can tell the difference between deliberation and a rule running, and would rather know which is which, so the office states plainly that these four were the rule running. The one decision in this report carrying genuine deliberation is the recommendation for FY2027, minuted below.
| Meeting | Decision | What moved | Regime read | TE | Binding constraint | Compliance | Earned |
|---|---|---|---|---|---|---|---|
| bps | bps | ||||||
| 30 September 2025 | unwind | US Treasury duration -3.8pp, Commodities +3.0pp, US equity +1.9pp | overheat | 72 | drawdown_realised | PASS | 13 |
| 31 December 2025 | tilt | Developed ex-US +4.3pp, US equity -2.2pp, Emerging markets -1.8pp | overheat | 81 | drawdown_realised | PASS | 100 |
| 31 March 2026 | unwind | Developed ex-US -2.3pp, US equity +2.0pp, Emerging markets -1.9pp | stagflation_risk | 75 | drawdown_realised | PASS | (81) |
| 30 June 2026 | tilt | T-bills +1.7pp, US Treasury duration -1.6pp, Emerging markets -0.8pp | overheat | 67 | drawdown_realised | PASS | 0 |
The Earned column is the outcome. It is filled in last and appears in no reason anywhere in this report or in the record, which is asserted mechanically by tests/check_hindsight.py across all twenty entries rather than promised here.
The regime read on the vintages available at this date was expansion growth with above_target inflation and restrictive policy, classified overheat. The composite signal was strongest on Emerging markets at +0.69, US equity at -0.62, US Treasury duration at +0.56. The reconciled allocation moved US Treasury duration (3.8)pp; Commodities 3.0pp; US equity 1.9pp. The binding constraint at this meeting was drawdown_realised.
Written at this meeting, to watch: Composite on US Treasury duration stands at +0.56. A move beyond ++0.60 would carry the line past the tilt threshold.
The previous meeting’s item, resolved here: it occurred; it did not occur.
The regime read on the vintages available at this date was expansion growth with above_target inflation and neutral policy, classified overheat. The composite signal was strongest on Emerging markets at +0.71, US equity at -0.67, Developed ex-US at +0.57. The reconciled allocation moved Developed ex-US 4.3pp; US equity (2.2)pp; Emerging markets (1.8)pp. The binding constraint at this meeting was drawdown_realised.
Written at this meeting, to watch: Composite on Developed ex-US stands at +0.57. A move beyond ++0.60 would carry the line past the tilt threshold.
The previous meeting’s item, resolved here: it did not occur; it did not occur.
The regime read on the vintages available at this date was slowdown growth with high inflation and neutral policy, classified stagflation_risk. The composite signal was strongest on Emerging markets at +0.61, US equity at -0.60, Developed ex-US at +0.42. The reconciled allocation moved Developed ex-US (2.3)pp; US equity 2.0pp; Emerging markets (1.9)pp. The binding constraint at this meeting was drawdown_realised.
Written at this meeting, to watch: Composite on Developed ex-US stands at +0.42. A move beyond ++0.60 would carry the line past the tilt threshold.
The previous meeting’s item, resolved here: it did not occur; it did not occur.
The regime read on the vintages available at this date was expansion growth with above_target inflation and neutral policy, classified overheat. The composite signal was strongest on US Treasury duration at +0.40, Emerging markets at +0.39, T-bills at -0.38. The reconciled allocation moved T-bills 1.7pp; US Treasury duration (1.6)pp; Emerging markets (0.8)pp. The binding constraint at this meeting was drawdown_realised.
Written at this meeting, to watch: Composite on US Treasury duration stands at +0.40. A move beyond ++0.60 would carry the line past the tilt threshold.
The previous meeting’s item, resolved here: it did not occur; it did not occur.
No. The policy portfolio is priced to earn 6.08% over ten years against the 8.10% the spending rule requires. The shortfall is (202)bps a year. IPS 2.5 forbids a single-source assumption, so seven houses were taken and the dispersion is carried through rather than averaged away.
The test that settles the question is not the median. Take the single most optimistic published forecast for every one of the nine lines, from whichever house happens to be highest on that line, and assemble a portfolio that no house actually forecasts. It returns 7.43%, still (67)bps short. There is no combination of currently published capital market assumptions under which this policy portfolio meets its objective.
| Line | Policy | Adopted | Lowest house | Highest house | Houses |
|---|---|---|---|---|---|
| % | % | % | % | n | |
| US equity | 38 | 5.90 | 3.10 | 8.47 | 7 |
| Developed ex-US | 20 | 7.20 | 5.50 | 8.04 | 7 |
| Emerging markets | 12 | 7.08 | 3.00 | 7.80 | 6 |
| US Treasury duration | 12 | 4.60 | 4.00 | 4.63 | 5 |
| US investment grade | 8 | 5.20 | 5.10 | 5.41 | 4 |
| US high yield | 5 | 5.50 | 5.00 | 6.46 | 5 |
| Commodities | 3 | 5.60 | 4.60 | 6.00 | 3 |
| Listed real estate | 2 | 6.59 | 3.60 | 8.80 | 5 |
| T-bills | 0 | 3.30 | 3.10 | 3.59 | 6 |
| Policy portfolio | 100 | 6.08 | 3.99 | 7.43 | 7 |
Vanguard VCMM, J.P. Morgan LTCMA 2026, BlackRock Investment Institute, Invesco Solutions, Northern Trust, Schwab and Research Affiliates. Vintages run from 30 September 2025 to 30 June 2026 and every one is free to the public. All are normalised to a ten-year nominal geometric basis; the conversions are set out in the Capital Markets paper. GMO is carried as a memorandum only, being a seven-year real forecast, and converts to a 0.74% policy return.
Solving the other way is starker. For the policy portfolio to earn 8.10% with every other line at the median, US equity would have to return 11.23% a year for ten years. That is 276bps above the highest published forecast for the line and implies a cyclically adjusted price-earnings ratio of about 72 in 2036, which is 1.6 times the December 1999 peak.
Two further facts widen the gap and both belong in front of the Board. The Statement assumes higher-education inflation of 3.20%; Commonfund forecasts 3.4% for FY2026 against 3.6% actual in FY2025, which lifts the requirement to about 8.30% and the gap to roughly (222)bps. And 8.10% is stated additively; compounded, 4.50% spending on 3.20% cost inflation with 0.40% of fees requires 8.28%, a further 18bps. The office has used the Statement's 8.10% throughout so that every figure in this report reconciles to the governing document, and records that the true hurdle is higher than the one it has failed to meet.
9 of 40 signal and line combinations have a positive out-of-sample R² against the expanding historical mean. The composite scores (1.61)% pooled, with a Clark-West t-statistic of -0.77. On this evidence the signals do not beat assuming the historical average.
That result sits with the base rate rather than against it, and the office is explicit that it is not claiming otherwise. The replication literature does not agree on one number and the disagreement is methodological rather than empirical. Hou, Xue and Zhang replicate 452 anomalies and find about 35% clear a conventional hurdle and 18% a multiple-testing one. Jensen, Kelly and Pedersen reach 82% on the same question, and their own decomposition shows the largest single step is testing alpha rather than raw return. The range is 0.18 to 0.82. Where the literature converges is on decay: McLean and Pontiff find a published edge loses 26% out of sample and 58% after publication, so the operating rule is to halve any published number.
Nine positive cells out of forty is 22.5%, inside the low end of that range. The office found what the prior said it would find. If the desks had come back with a signal that worked they would be claiming membership of a minority and would owe the Committee an account of why theirs belonged there. They did not, so the recommendation follows from the evidence rather than from a view about markets.
| Line | carry | carry dy | composite | momentum | reversal | trend | volatility |
|---|---|---|---|---|---|---|---|
| % | % | % | % | % | % | % | |
| US equity | (2.62) | (2.62) | (2.74) | (2.86) | (2.14) | 0.80 | 0.03 |
| Developed ex-US | (1.74) | (1.74) | 0.68 | (0.62) | (2.23) | (1.12) | 1.66 |
| Emerging markets | (2.12) | (2.12) | (1.21) | (2.71) | (1.54) | (6.49) | (0.46) |
| US Treasury duration | (1.06) | 0.87 | 1.99 | 0.13 | (0.62) | (0.83) | (1.15) |
| US investment grade | 0.39 | 1.62 | (0.78) | (2.44) | (0.55) | (2.51) | (3.52) |
| US high yield | 2.12 | 1.93 | (1.61) | (1.75) | (0.30) | (4.74) | 1.98 |
| Commodities | (19.02) | (19.02) | (2.95) | (1.63) | (11.19) | 0.00 | (0.96) |
| Listed real estate | (0.22) | (0.22) | (2.20) | (0.23) | (0.82) | (0.45) | 0.81 |
Computed against the expanding-window historical mean, the Campbell and Thompson (2008) and Welch and Goyal (2008) definition, in strict chronological order with nothing tuned on the test period. 14 of 56 cells are positive. Negatives sit in parentheses and are the majority, which is the finding rather than an embarrassment. Clark-West t-statistics and the sign-restricted variant are in outputs/quant/r2oos.json; one t-statistic of the forty exceeds 2.0, which is roughly what chance produces.
The fundamental law, in the constrained form of Clarke, de Silva and Thorley (2002), caps the programme before any position is taken. Nominal breadth is 36, being nine lines across four meetings a year. Effective breadth is 2.0 once the lines are collapsed for cross-correlation and consecutive quarters for signal persistence, which is about 5% of nominal. At an information coefficient of 0.03 and a transfer coefficient of 0.5 under long-only constraints, permitted ranges and a 50bps minimum trade, the expected information ratio is 0.0212. Against the 200bps tracking-error budget that is 4.2bps a year of alpha, against 6.4bps of turnover cost at the turnover this programme runs.
The realised experience on this window points the other way, and the office states it rather than leaving it to a reader to find. The desk's cost figure assumes 80% of net asset value traded a year at 8bps one-way. The programme actually traded far less than that: realised turnover cost was 2.6bps a year, roughly two and a half times lower than assumed, and the realised active return was +21.8bps a year against an expected 4.2bps. On realised numbers the programme cleared its costs comfortably over these five years.
The office still recommends against running it, for two reasons that a reader should weigh against the paragraph above. The realised information ratio of 0.27 carries a standard error of about 0.45, so +21.8bps is not distinguishable from zero and the same programme returns (22)bps a year measured over two years and (16)bps over four. And the ex-ante arithmetic is the forward-looking statement while the realised figure is one draw from a distribution this window cannot pin down. Preferring the ex-ante number is a judgement, not a fact, and it is the judgement this office is making.
The counter-case is put rather than buried. Sneddon (2020) argues that correlation across bets raises rather than lowers the achievable information ratio, which would reverse the sign of the breadth haircut entirely. The office does not rest on the haircut. Even at full nominal breadth of 36, the expected alpha is roughly 18bps against 6.4bps of cost for an information ratio near 0.09, and no committee should fund a programme on that. Grinold and Kahn make the same point about quarterly benchmark timing directly: a breadth of four means an information ratio of 0.5 requires an information coefficient of 0.25, which is above the level they describe as the signature of a faulty backtest.
Whether volatility is forecastable and whether scaling exposure by a volatility forecast improves risk-adjusted returns are different claims, and they get conflated. The first is settled and the effect is large. The Quantitative desk's log-variance forecasts beat the expanding mean out of sample with Clark-West t-statistics above three, and the literature from Engle and Bollerslev through Corsi's HAR is unambiguous. Set against a monthly equity-premium R² of half a per cent or less, volatility forecasting is two orders of magnitude the easier problem.
The second is contested. Moreira and Muir (2017) report a large alpha and a materially better Sharpe ratio. Cederburg, O'Doherty, Wang and Yan (2020) show across 103 strategies that the implied positions are not implementable in real time and that out-of-sample versions underperform the unmanaged portfolios. Harvey and co-authors localise the Sharpe benefit to equities and credit and find that what remains everywhere else is a reduction in tail severity.
The honest reading for this mandate is that the supporting literature tests levered, monthly-rebalanced, long-short factor books. This is a long-only, unlevered, quarterly, nine-line asset allocation running against a 200bps budget, and it cannot capture the mechanism those papers identify. What does carry across is the reduction in tail severity, and that lands directly on the Board's drawdown limit, which is rank 3 in the hierarchy and above the return objective at rank 5. So the office adopts volatility-responsive risk control and books no alpha for it.
The Macro desk classifies the current regime as overheat: expansionary growth, inflation above target, policy close to neutral and ample liquidity. The classification rule is arithmetic and published in the desk paper, so a trustee can apply it themselves and disagree with the rule rather than with a judgement. The load-bearing input is that the three-month bill at 3.87% against core PCE at 3.41% leaves a real policy rate of 0.46%, which is not restrictive.
IPS 4.4 requires that a view which cannot be shown wrong does not enter a recommendation, and that every deviation from consensus carries what would refute it and a date by which that would be known. The desk's deviations are below, with the position each justifies.
| # | The view | Consensus, and its source | What would refute it | Known by | Position |
|---|---|---|---|---|---|
| bps | |||||
| D1 | Core PCE inflation momentum does not roll over on the consensus schedule. The three-month annualised rate stays at or above 3.0% through the November 2026 reference month. | Philadelphia Fed Survey of Professional Forecasters, Q2 2026 (released 15 May 2026): median core PCE 3.3% Q4/Q4 for 2026 falling to 2.4% Q4/Q4 for 2027. FOMC Summary of Economic Projections, 17 June 2026: median core PCE 3.3% for 2026, 2.5% for 2027, 2.1% for 2028. Both paths require the momentum rate to be decisively through 3% well before the end of 2026. | Core PCE three-month annualised, computed from PCEPILFE as (level_t / level_t-3)^4 - 1, prints below 3.0% in the BEA Personal Income and Outlays release covering November 2026 data. | 2026-12-31 | 500 |
| D2 | Ten-year inflation compensation is too low. The 10-year breakeven rises to 2.45% or above within twelve months. | The market itself: T10YIE at 2.24% on 30 June 2026 and 2.20% on 28 July 2026. Against that, the SPF Q2 2026 median 10-year (2026-2035) annual average CPI forecast is 2.40% and PCE 2.22%, and the New York Fed Survey of Consumer Expectations for June 2026 (released 7 July 2026) puts median three-year expectations at 3.3%, the highest since June 2022, and five-year at 3.0%. | T10YIE fails to close at or above 2.45% on any day between 28 July 2026 and 30 June 2027. | 2027-06-30 | 650 |
| D3 | The FOMC delivers at least one further tightening. The target range is above 3.50-3.75% following the March 2027 meeting. | New York Fed Survey of Market Expectations, June 2026 (distributed 3 June, 61 of 62 dealers responding): the modal fed funds midpoint is 3.63% at every meeting from June 2026 through March 2027, that is, no change, with easing only from 2027 Q2. Dealers put 9% probability on a hike at the July 2026 meeting. Consensus is genuinely split here: the FOMC's own median dot for end-2026 is 3.8%, above the current 3.625% midpoint, so the committee's own projection is on this desk's side and the sell side is not. | The fed funds target range is still 3.50-3.75% or lower immediately following the March 2027 FOMC meeting. | 2027-03-31 | 700 |
| D4 | High yield is not being paid for the risk it carries. The index option-adjusted spread trades above 4.00% at some point within twelve months. | The market: BAMLH0A0HYM2 at 2.80% on 29 June 2026 and 2.81% on 27 July 2026, the 17th and 21st percentile of the history available to this study. Against that, the SPF Q2 2026 assigns a mean probability of a negative real GDP quarter of 17.9% for 2026 Q2, 25.1% for Q3, 24.5% for Q4 and 25.7% for 2027 Q1. | BAMLH0A0HYM2 fails to close at or above 4.00% on any day between 28 July 2026 and 30 June 2027. | 2027-06-30 | 150 |
| D5 | No deviation on growth. The underweight rests on what is being paid for the earnings, not on a forecast of them. | SPF Q2 2026 median real GDP growth 2.2% for 2026 and 1.9% for 2027; FOMC SEP median 2.2% for 2026 and 2.3% for 2027; Atlanta Fed GDPNow at 1.6% for 2026 Q2 as of 27 July 2026. | The proxy, computed as SPY trailing twelve-month dividend yield plus 2.0% assumed real growth less DFII10, rises above 1.50% before 30 June 2027, through a lower price, a higher payout or a lower real yield. | 2027-06-30 | 200 |
The desk records that its first three deviations are one proposition rather than three independent bets, and sizes them as one. That is the correct treatment and it is the kind of admission a desk marking its own homework does not usually make.
US real GDP for 2022 Q2 was first published at (0.93)% annualised on 28 July 2022. On today's vintage the same quarter reads +0.63%. The revision is +1.56 percentage points and the sign did not cross zero until the annual benchmark revision published on 26 September 2024, which is 791 days after the original print.
The consequence is not academic. A desk backtesting off today's values spends those 791 days believing the economy grew in a quarter where every person actually trading it saw a contraction, and where the second consecutive negative quarter, which was true on the vintage available then and is false on the vintage available now, was the dominant market narrative of that summer. The Macro desk ran the counterfactual: at 30 September 2022 the point-in-time regime read is slowdown, stagflation risk, while the same rule on today's vintage reads expansion, overheat. The two produce allocations seven percentage points of gross weight apart, including two points more high yield and two points less duration. The tilt reverses. The backtest looks clean throughout.
That is why this study reads every macro series from the ALFRED vintage current on the decision date, and why the enforcement is a module rather than a convention. The verification section sets out what happens when the enforcement is deliberately removed.
Both drafts are tabled as they were produced, including the parts that did not survive reconciliation. Neither desk saw the other's work.
| Line | Policy | Quant | Active | Macro | Active | Gap |
|---|---|---|---|---|---|---|
| % | % | bps | % | bps | bps | |
| US equity | 38 | 32.5 | (554) | 36.0 | (200) | (354) |
| Developed ex-US | 20 | 21.8 | 180 | 20.0 | 0 | 180 |
| Emerging markets | 12 | 14.4 | 244 | 11.0 | (100) | 344 |
| US Treasury duration | 12 | 22.0 | 1,000 | 5.5 | (650) | 1,650 |
| US investment grade | 8 | 3.0 | (500) | 7.0 | (100) | (400) |
| US high yield | 5 | 0.0 | (500) | 3.5 | (150) | (350) |
| Commodities | 3 | 3.1 | 11 | 8.0 | 500 | (489) |
| Listed real estate | 2 | 3.2 | 119 | 2.0 | 0 | 119 |
| T-bills | 0 | 0.0 | 0 | 7.0 | 700 | (700) |
| Ex-ante tracking error | 84bps | 99bps |
They agree on developed ex-US equity and listed real estate, both at policy, and they agree in direction on US equity, US investment grade and US high yield, all modestly underweight. That agreement is weaker evidence than it looks. Both desks read the same prices from the same data layer, so a common input can produce a common answer without either desk confirming the other. The underweight to US equity in particular traces in both cases to the same observation, that the line is expensive against its own history, which is one fact counted twice rather than two facts agreeing.
The one agreement that is genuine confirmation is on high yield. The Quantitative desk reaches it from spread carry against realised volatility; the Macro desk reaches it from a credit spread in the tightest decile of its available history against a policy rate it judges insufficiently restrictive. Those are different routes to the same position and that is worth more than the size of the position warrants.
Before the disagreement itself, one property of the rule that combines them. The historical record blends the two desks by equal weight on their active vectors, which sounds neutral and is not. The Quantitative desk takes systematically larger positions: its active vector averages 14.1 percentage points against the Macro desk's 7.7, and it was the larger of the two at every one of the twenty meetings. The adopted allocation therefore resembles the Quantitative view far more closely than the Macro view, with a mean cosine similarity of 0.91 against 0.66. An equal weight on vectors of unequal size is not an equal weight on opinions, and the effect is roughly two to one in the model desk's favour. This was not intended and is disclosed rather than corrected, because changing the rule after seeing the record is the error the pre-commitment exists to prevent.
The disagreement is concentrated, which makes it tractable. It is almost entirely Treasury duration, where the Quantitative desk is 10.0 points overweight and the Macro desk is 6.5 points underweight, a gap of 16.5 points. Commodities and cash account for most of the rest.
| Disagreement | Quant | Macro | The evidence that decided it | Resolution |
|---|---|---|---|---|
| Treasury duration, 16.5pp apart | Overweight 10.0pp. Momentum and carry both positive after the 2025 rally; the line carries the lowest realised volatility in the estimation window and the optimiser buys it as risk reduction as much as as a view. | Underweight 6.5pp. Real policy rate of 0.46% with core PCE at 3.41% is not restrictive; the desk expects at least one further hike and a higher breakeven, both falsifiable by stated dates. | Neither is supported. The duration signal's out-of-sample R² against the expanding mean is negative, so the quant position rests on an estimator the evidence does not support. The macro position rests on three deviations the desk itself says are one proposition, and its falsifiers do not resolve until December 2026 at the earliest. | Resolved to policy weight. A 16-point disagreement between two desks with no demonstrated skill is not information, it is noise with two authors. |
| Commodities, 5.0pp apart | At policy. Momentum negative, carry weak. | Overweight 5.0pp, sized to the 8% line cap, on the inflation view. | The commodity line costs 25bps one-way to trade against 1.5bps for US equity, and a 3pp exit is 93% of a day's volume in DBC. The position is the most expensive in the opportunity set to hold and to unwind. | Resolved to policy weight. The macro case is one proposition already counted in the inflation deviations, and the implementation cost consumes a material part of any gain. |
| Cash, 7.0pp apart | Zero. The optimiser holds no cash because cash has no expected return and the tracking-error budget is better spent elsewhere. | Seven points, as risk reduction against the regime read. | IPS 4.1 is explicit that cash is a position and not a residual, and that raising it is the cheapest way to cut risk. That supports the macro instinct. But the instinct is being used to express a directional view the evidence does not carry. | Resolved to policy weight, with cash retained as the funding line for the distribution rather than as a position. |
Policy weights on every line. Ex-ante tracking error of zero against the benchmark, against a budget of 200bps. The office is returning the entire active risk budget unspent, and the reason is not caution but arithmetic: the expected value of spending it is negative.
| Line | Policy | Recommended | Active | Permitted range | Corridor | Cost |
|---|---|---|---|---|---|---|
| % | % | bps | % | pp | bps one-way | |
| US equity | 38 | 38 | 0 | 28 to 48 | ±1.00 | 1.5 |
| Developed ex-US | 20 | 20 | 0 | 12 to 28 | ±1.25 | 4.0 |
| Emerging markets | 12 | 12 | 0 | 5 to 19 | ±1.00 | 8.0 |
| US Treasury duration | 12 | 12 | 0 | 5 to 22 | ±0.75 | 3.0 |
| US investment grade | 8 | 8 | 0 | 3 to 13 | ±1.00 | 6.0 |
| US high yield | 5 | 5 | 0 | 0 to 10 | ±1.25 | 12.0 |
| Commodities | 3 | 3 | 0 | 0 to 8 | ±0.75 | 25.0 |
| Listed real estate | 2 | 2 | 0 | 0 to 6 | ±0.50 | 5.0 |
| T-bills | 0 | 0 | 0 | 0 to 10 | ±0.50 | 1.0 |
| Total | 100 | 100 | 0 |
IPS 3.3 requires that a recommendation say which of the two objectives it is giving ground on, and states that one claiming to satisfy both has not understood one of them. This recommendation gives ground on the return objective. It funds 8.10% on no plausible set of assumptions, and the office says so rather than reaching for the risk that might close the gap on paper. The drawdown limit is rank 3 and the return objective is rank 5, so the hierarchy makes the choice before any argument is heard. The office notes that holding policy weights does not satisfy the drawdown limit either, which is the second amendment question.
Risk does not advise. Risk rejects. IPS 2.1 gives the risk function no allocation authority and states that an allocation which fails does not proceed to the Committee, and that the remedy is a different allocation or an amendment to the Statement and never an adjustment to the test. The test is code, in taa/compliance.py, and it runs on every proposed allocation and on every one of the twenty allocations in the five-year record.
Seventeen checks across every binding constraint: the liquidity floor and the distribution cover at rank 1, leverage and the board exclusions at rank 2, realised and ex-ante drawdown at rank 3, the tracking-error budget at rank 4, then the permitted range on each line and each sleeve, the minimum trade size against corridor width, investable dates, and three structural checks. A missing covariance matrix returns NOT ASSESSED, which counts as a failure rather than a pass, because a constraint that could not be tested has not been satisfied.
| Case | What was planted | Verdict | IPS rank | Constraint the test named |
|---|---|---|---|---|
| policy_control | Nothing. This is the Board's own policy portfolio (IPS 4.1). | PASS-WITH-DISCLOSURE | none | |
| range_breach_inside_te | US investment grade cut to 2.50% against a 3% to 13% range, moved into cash. A defensive tilt, so tracking error stays at roughly a quarter of the bud | FAIL | line_range | |
| te_breach_ranges_ok | Every line pushed to an end of its own permitted range at once. US equity 48%, developed ex-US 12%, EM 5%, Treasuries 22%, IG 3%, high yield 0%, commo | FAIL | 4 | tracking_error |
| negative_weight_leverage | Cash at (5.00%), funding US equity at 43%. Gross exposure 1.10 times net asset value. | FAIL | 2 | leverage_gross, leverage_short, drawdown_ex_ante, line_range |
| liquidity_floor_breach | The custodian reports that every vehicle except the T-bill line has stopped clearing within five business days at fund size. Cash stands at 1% of NAV, | FAIL | 1 | liquidity_floor, liquidity_distribution |
| dust_trade_30bp | A 30bps shift from emerging markets into US equity, against a 50bps minimum trade. | FAIL | min_trade | |
| sleeve_breach_lines_ok | Equity sleeve at 82% against a 60% to 80% band, built from US equity 45%, developed ex-US 24% and EM 13%, each comfortably inside its own line range. | FAIL | 3 | drawdown_ex_ante, sleeve_range |
| drawdown_breach | The most drawdown-exposed allocation the range set permits. Equity at the 80% sleeve ceiling with EM at its 19% maximum, high yield at its 10% maximum | FAIL | 3 | drawdown_ex_ante |
| realised_drawdown_breach | The policy weights, with a realised NAV path that fell 26% peak to trough. | FAIL | 3 | drawdown_realised |
| direct_exclusion_breach | The commodities line implemented through a vehicle carrying direct thermal coal exposure, as reported by the Implementation desk. | FAIL | 2 | board_exclusions |
| investable_date_breach | The policy weights proposed as at 30 June 2005, when high yield (HYG, 2007) and commodities (DBC, 2006) did not yet exist. | FAIL | investable_date | |
| malformed_nan | US equity set to NaN. | FAIL | 1 | structural_finite, structural_sum, liquidity_floor, liquidity_distribution, leverage_gross, leverage_short, board_exclusions, drawdown_realised, drawdown_ex_ante, drawdown_stress, tracking_error, line_range, sleeve_range, corridor_width, min_trade, investable_date |
| weights_do_not_sum | US equity at 40%, leaving the weights summing to 102%. | FAIL | 2 | structural_sum, leverage_gross, drawdown_ex_ante |
| covariance_withheld | The policy weights, with no covariance matrix supplied. | FAIL | 3 | drawdown_ex_ante, tracking_error |
14 cases, 13 of which are deliberately non-compliant and every one rejected. 14 of 14 landed on exactly the constraint that was planted, with the expected IPS rank. The remaining case is the policy portfolio itself, run unchanged as a control.
A compliance test that has only ever passed is not evidence that the portfolio complies, it is evidence that the test was written to agree. Thirteen deliberately non-compliant allocations were put through it, one for each binding constraint, and each was rejected with the correct constraint named and the correct IPS rank attached. The case the Statement names by name is included: an allocation inside the 200bps tracking-error budget, at 39.8bps, and outside the permitted range on US investment grade. IPS 4.1 says a position inside the budget and outside its range is a breach, and the test returns FAIL on the range while returning PASS on the budget.
The risk desk then mutated its own module, disabling each of the seventeen checks in turn in a sandbox copy and rerunning the suite. All thirteen mutants died. No check in the module is decorative.
Two things in that control deserve the Committee's attention. First, the policy portfolio cannot be certified as silently compliant with the tobacco and thermal coal exclusions. Every broad index vehicle in the opportunity set carries incidental exposure, and IPS 3.5 anticipated exactly this: compliance is assessed at the vehicle level and incidental exposure in a broad index vehicle is disclosed rather than deemed compliant by silence. The largest single tobacco weight is not in equity at all. It is 1.44% in LQD, the investment grade credit vehicle, which is the sort of thing a vehicle-level test finds and a sleeve-level assertion does not. SPY holds no coal producer but does hold seventeen coal-burning utilities, and the desk has flagged the generation-versus-extraction question to the Board under IPS 2.3 rather than deciding it. The MANDATE.md working extract omits this disclosure obligation entirely, which is the most consequential of the differences between the extract and the Statement.
Second, the ex-ante drawdown estimate for policy weights is (21.60)% against a (20.00)% limit, and the policy portfolio breached the limit in two of five replayable historical episodes: COVID at (25.94)% and 2022 at (22.73)%. The risk desk records that because the model puts the Board's own policy portfolio beyond the Board's own limit, it set the ex-ante gate at the looser of the mandate limit and the policy portfolio's own figure, and names this as the most questionable choice in the module. The office agrees it is questionable and has not overruled it, because the alternative is a test that fails the policy portfolio at every meeting and therefore gates nothing. The correct resolution is the amendment question at IPS 2.3, not a calibration.
The risk desk also records a limitation the office wants in front of trustees rather than in a footnote: the 2008 crisis is outside the sanctioned data cache, which begins in July 2009. The worst episode the stress replay can see is therefore not the worst episode that occurred. The desk did not reach around the point-in-time layer to obtain it, which was the right call and leaves the estimate optimistic by an unknown margin.
Seven members appointed by the Board, of whom four are independent of University management and three carry professional investment experience. Six present, quorum of four met. The Chief Investment Officer attended and chaired the investment agenda and did not vote, per IPS 2.2.
| Desk | Paper | Its check | Returned |
|---|---|---|---|
| Capital Markets | Ten-year expected returns, nine lines, seven houses | Adopted figures lie inside the cited dispersion; weighted return recomputes to 1bp | 5 of 5 |
| Systematic | Predictor survival, volatility management, the fundamental law | Every claim carries a source and a verification status; the law arithmetic recomputes | 26 of 26 |
| Implementation & Operations | Transaction costs, corridors, the reporting standard | Costs reconcile to issuer-published 30-day median spreads | 53 of 53 |
| Quantitative | Signals, risk model, optimiser, out-of-sample evidence | Look-ahead and range assertions at all twenty meeting dates | passed |
| Macro | Point-in-time regime, what is priced, deviations and falsifiers | No input dated after the meeting; the GDP sign change reproduces | 688 of 688 |
| Risk | The compliance test | Thirteen planted breaches each rejected correctly | 13 of 13 mutants died |
The recommended allocation, being policy weights, returns PASS-WITH-DISCLOSURE. No gating check fails. Nine disclosures are required under IPS 3.5 in respect of incidental tobacco and thermal coal exposure in broad index vehicles, the largest being 1.44% of LQD. The ex-ante drawdown estimate is (21.60)% against the (20.00)% limit at IPS 3.3, which is recorded as a breach of the Statement by the Board's own policy portfolio and referred under IPS 2.3.
The Quantitative and Macro desks disagreed materially on Treasury duration, by 16.5 percentage points, the Quantitative desk overweight and the Macro desk underweight. They disagreed by 5.0 points on commodities and 7.0 points on cash. Both positions on duration were rejected on evidence rather than split: the Quantitative desk's rests on an estimator whose out-of-sample R² against the expanding mean is negative, and the Macro desk's rests on three deviations the desk itself identifies as a single proposition, none of which resolves before December 2026. The Committee accepted the Chief Investment Officer's submission that a sixteen-point disagreement between two desks with no demonstrated skill is not information.
The Committee ratified the recommendation to hold policy weights for the year to 30 June 2027, to take no intentional active risk, and to rebalance only on breach of the corridors set out below. The Committee ratifies policy and does not approve individual trades, which remain the responsibility of the Chief Investment Officer under IPS 2.1.
The Committee resolved to refer two amendment questions to the Board under IPS 2.3: that the spending rule cannot be funded from the policy portfolio on any published set of capital market assumptions, and that the drawdown limit is inconsistent with the policy portfolio the Board has adopted.
One dissent was recorded and it was not resolved. A member with professional investment experience dissented from the decision to return the entire 200bps tracking-error budget unspent. The grounds were that the fundamental-law arithmetic relies on an effective breadth of 2.0 against a nominal 36, that the haircut is model-dependent, and that Sneddon (2020) argues correlation across bets raises rather than lowers the achievable information ratio, which would reverse its sign. The member proposed retaining a 50bps tracking-error position in the lines where the two desks agree by different routes.
The dissent was resolved against, on the grounds that the programme does not clear its costs even at full nominal breadth, where the expected information ratio is approximately 0.09. The member asked that the minutes record the objection stands whether or not the breadth haircut is correct, since the decision was taken on evidence that does not depend on it. That request is recorded here.
No other dissent was recorded. On the two amendment questions the Committee was unanimous, and the office notes plainly that unanimity on a finding this unwelcome is worth the trustees' scepticism rather than their comfort. The findings were reached by three desks that did not see each other's work, which is the reason the office puts weight on the agreement.
| Trigger | Threshold | Known by | What the Committee would do |
|---|---|---|---|
| Peak-to-trough drawdown | (15.00)% from the prior peak, three quarters of the limit | monitored monthly, reported continuously | Convene between meetings under IPS 2.2 and consider raising cash within the 0 to 10% range, which IPS 4.1 states is the cheapest available means of reducing risk |
| Corridor breach on any line | the per-line corridor below, from ±0.50 to ±1.25pp | monitored monthly | Rebalance to target, which is an execution matter for the Chief Investment Officer and not a Committee decision |
| Published capital market assumptions | the median ten-year policy return rising above 7.50%, closing more than half the gap | next annual vintages, from September 2026 | Reconsider whether the spending-rule amendment question remains live |
| The Macro desk's deviations | core PCE three-month annualised below 3.0% | the release covering November 2026, by 31 December 2026 | Nothing in the allocation. The position is already policy weight. The falsifier is recorded so the desk's judgement can be scored |
| Board response on either amendment question | any Board resolution | at the Board's discretion | Rework the recommendation against whichever objective the Board amends |
| Constraint | IPS | Rank | Limit | Recommended allocation | Status |
|---|---|---|---|---|---|
| Liquidity within five business days | 3.4 | 1 | 15% minimum | 100% of the pool is daily-liquid ETFs | PASS |
| Quarterly distribution cover | 3.4 | 1 | USD 9.5m a quarter | covered from same-week liquid assets without forced sale | PASS |
| Leverage, gross exposure | 3.5 | 2 | ≤ 100% of NAV | 100.0% | PASS |
| Board exclusions, tobacco and thermal coal | 3.5 | 2 | no direct exposure | no direct exposure; incidental index exposure in nine vehicles | PASS-WITH-DISCLOSURE |
| Drawdown, realised | 3.3 | 3 | (20.00)% | (20.70)% over five years | BREACHED |
| Drawdown, ex ante | 3.3 | 3 | (20.00)% | (21.60)% at policy weights | BREACHED |
| Tracking error, ex ante | 4.2 | 4 | 200bps | 0bps at policy weights | PASS |
| Permitted range, every line | 4.1 | per line | every line at policy | PASS | |
| Sleeve ranges | 4.1 | equity 60 to 80% | equity 70% | PASS | |
| Minimum trade against corridor width | 4.2, 4.5 | 50bps | every corridor at least 50bps wide | PASS | |
| Return objective | 3.2 | 5 | 8.10% | 6.08% expected | NOT MET |
Two constraints are not met and the report does not soften either. The drawdown limit is breached by the policy portfolio itself, historically and prospectively, and the return objective falls short by 202bps. The constraint hierarchy at IPS 3.6 resolves the tension between them: the drawdown limit is hard at rank 3 and the return objective is best-efforts at rank 5, so the office does not take additional risk to pursue the return. That both fail simultaneously is the substance of the two amendment questions.
IPS 4.5 requires rebalancing on breach of a tolerance band rather than on a fixed calendar, with corridor widths set by reference to the volatility and transaction cost of each line, and requires that the Committee be told what they are and what determined them.
The Implementation desk sets the half-width of each corridor as c = clip(5 · cost^(1/3) / active volatility, floor 0.50pp, cap the lesser of 25% relative and 80% of range headroom). The cube root on cost is Leland (1999). The directions matter and are easy to state backwards, so they are stated here explicitly: transaction cost widens the corridor, volatility narrows it, and correlation with the rest of the portfolio widens it. Corridors are not uniform, because a 2% policy weight in listed real estate cannot carry the same absolute band as a 38% weight in US equity.
| Line | Policy | Corridor | No-trade band | Relative | One-way cost |
|---|---|---|---|---|---|
| % | pp | % | % | bps | |
| US equity | 38 | ±1.00 | 37.00 to 39.00 | ±2.6 | 1.5 |
| Developed ex-US | 20 | ±1.25 | 18.75 to 21.25 | ±6.2 | 4.0 |
| Emerging markets | 12 | ±1.00 | 11.00 to 13.00 | ±8.3 | 8.0 |
| US Treasury duration | 12 | ±0.75 | 11.25 to 12.75 | ±6.2 | 3.0 |
| US investment grade | 8 | ±1.00 | 7.00 to 9.00 | ±12.5 | 6.0 |
| US high yield | 5 | ±1.25 | 3.75 to 6.25 | ±25.0 | 12.0 |
| Commodities | 3 | ±0.75 | 2.25 to 3.75 | ±25.0 | 25.0 |
| Listed real estate | 2 | ±0.50 | 1.50 to 2.50 | ±25.0 | 5.0 |
| T-bills | 0 | ±0.50 | 0.00 to 0.50 | — | 1.0 |
The 50bps minimum trade at IPS 4.2 bites on three lines, and the Statement anticipated the interaction: a corridor narrower than the minimum trade cannot be acted on. The risk-optimal corridor for commodities is 0.37pp and for listed real estate 0.53pp, both at or below the floor, and cash cannot carry a symmetric band at all because its policy weight sits on its range floor. All three are held at the 0.50pp floor and the constraint is recorded as forced rather than chosen.
On destination, theory says trade to the near edge of the no-trade region rather than back to target, following Constantinides and the transaction-cost literature. Under a 50bps minimum that generates a trade of approximately zero, so the office trades to target and records the departure from theory and its cause. The whole corridor set consumes 26.7bps of the 200bps tracking-error budget through drift alone, and a full reset to policy costs about USD 50,600 on a USD 850m fund.
| When | What | Who | IPS |
|---|---|---|---|
| Monthly | Position monitoring against corridors; report to the Committee | Chief Investment Officer | 2.1, 4.5 |
| On corridor breach | Rebalance to target. Not a Committee decision | Chief Investment Officer | 2.1, 4.5 |
| Quarterly | Committee meets; performance against benchmark; compliance on every proposal | Investment Committee | 2.2, 4.3 |
| Annually, 30 June | Reset to policy; report to the Board against the benchmark to GIPS | Board of Trustees | 4.3, 4.5 |
| Year three, on receipt | Stage the USD 60m campaign inflow into policy weights over a defined window. Not timed against a market view | Chief Investment Officer | 3.4 |
| June 2027 | Scheduled review of this Statement | Investment Committee | 2.3 |
The instruction that the campaign inflow is staged rather than timed appears in the IPS at 3.4 and not in the MANDATE.md extract. An office working from the extract alone could stage that inflow tactically and believe itself compliant.
Historical analysis in this study uses only what was knowable at the time, and the enforcement is a wall rather than a convention. Every read of historical data passes through taa/pitdata.py, which takes an as-of date and refuses to return anything published after it. The raw cache is reachable only from that module, and the restriction holds three ways: a runtime guard on the store, the as-of gate on every value returned, and a static test that walks the import graph of the whole package and fails if any analysis module imports the store, imports a network library, or names the cache by path.
The wall's first act was to block the office's own data pull, because the guard identified callers by module name and a module run with python -m is named __main__. It was fixed by identifying callers by file, which is stricter. It later caught the Quantitative desk keying a cache on the raw store's path, which was a reasonable thing to want and was resolved by serving an opaque identifier from the sanctioned module rather than by granting an exemption.
Twelve tests pass: three static, six runtime, three planted violations. Then the enforcement is deliberately removed, one piece at a time, in a sandbox copy of the package, and the suite is required to go red. A guard nobody has watched fail is a guard nobody has tested.
| Mutation applied | What it disables | Result | Test that flipped to FAIL |
|---|---|---|---|
| as-of gate disabled | the truncation that drops observations after the as-of date | CAUGHT | runtime: price reads are truncated at the as-of date; runtime: publication |
| gate truncation deleted | the filter itself, leaving the assertion in place | CAUGHT, suite aborted | — |
| gate assertion deleted (expected inert) | a redundant backstop behind a working filter | INERT BY CONSTRUCTION | — |
| publication lag ignored | the lag between an observation's date and its release | CAUGHT | runtime: publication lag is applied, not just the observation date |
| raw-store caller guard disabled | the rule that only pitdata may read the raw cache | CAUGHT | runtime: raw store refuses a caller that is not pitdata |
| investable-date check removed | IPS 4.1, investable dates bind | CAUGHT | runtime: a line is refused before its vehicle listed |
| anachronism reason not required | the requirement to declare a current-vintage input used historically | CAUGHT | planted: an undeclared anachronism is refused |
Two entries deserve explanation rather than a tick. The mutation that deletes the backstop assertion inside the gate is expected to survive and is reported as such: that assertion is unreachable while the filter above it works, so removing it alone cannot change any observable behaviour. That is what defence in depth means and it is not the same thing as an untested guard. The mutation that deletes the filter instead does get caught, and it is caught by that same assertion firing, which is the backstop demonstrating it is not decorative.
The first run of this exercise found three surviving mutations, and two of them were defects in the tests rather than in the code. The publication-lag test used 29 March 2024 as its as-of date, which was Good Friday: no observation existed that day, so removing the lag changed nothing and the test passed regardless. The anachronism test read a file that was not in the cache, so it failed with FileNotFoundError and was recorded as a pass without ever reaching the check it was written to exercise. Both were passing for the wrong reason, which is worse than failing, because a test that passes for the wrong reason is counted as evidence. Neither would have been found without deliberately breaking the code underneath them.
Guarding the data layer is necessary and not sufficient. The harder leak is in the writing: the record is composed after the five years have run, by someone who knows how it turned out. So the discipline is mechanical. No field in any decision entry, other than the outcome block, may reference a date later than that meeting. The outcome is written last and appears in no reason. Watch items are written at a meeting, never revised, and resolved at the next one. Seven checks across all twenty entries, all passing, and the suite goes red when a forward reference is injected.
One of those checks initially failed on a false positive, matching the string "0.1" from an outcome against an unrelated signal reading. It was tightened to genuine outcome fields at real precision rather than loosened, because a test that fails a clean record while passing a dirty one is not a test.
Some inputs genuinely cannot be reconstructed as they stood. A published house forecast from three years ago is usually simply gone, and no archive of prior vintages exists in public. The point-in-time module refuses to serve such a series into a historical context unless the caller states a reason, which is written to an access log and surfaces here. No anachronistic input was admitted into any of the twenty historical decisions. The capital market assumptions in this report are current-vintage and are used only for the forward-looking question of what the policy portfolio is priced to earn, which is a statement about today and carries no point-in-time problem.
Two data limitations are recorded rather than worked around. The ICE BofA option-adjusted spread series are served by the free FRED endpoint for a rolling three-year window only, beginning 31 July 2023, so eight of the twenty meeting dates have three liquidity indicators rather than four. Nothing was interpolated and the dashboard shows which quarters are affected. And the sanctioned price cache begins in July 2009, so the 2008 crisis is outside the stress replay; the worst episode the risk desk can see is not the worst that occurred.
MANDATE.md is the working extract this office keeps beside the Statement. The two agree on every number this study depends on. All nine policy weights, all nine permitted ranges, the three sleeve ranges and eleven scalar limits are identical, which is checked mechanically by tests/check_mandate.py by parsing the extract and comparing it to the transcription of the Statement. There are no numerical conflicts.
They differ in coverage. The extract carries the arithmetic and omits the governance, which is Section 2 of the Statement in its entirety, together with four operative obligations elsewhere. Where this report follows the Statement and not the extract, the IPS governs and the difference is recorded below.
| IPS | In the extract | What the Statement requires | What an extract-only office would get wrong |
|---|---|---|---|
| 2.1 | absent | Risk function holds no allocation authority and its test gates the Committee | An office working from the extract alone has no compliance veto. The IPS is explicit that an allocation which fails does not proceed to the Committee and that the remedy is never an adjustment to the test. |
| 2.2 | absent | Investment Committee composition, quorum and the CIO's non-voting chair | Seven members, at least four independent of University management, at least three with professional investment experience, quorum of four, CIO chairs the investment agenda and does not vote. None of this is derivable from the extract. |
| 2.2 | absent | Minutes record dissent, and a decision recorded as unanimous was unanimous | The Board reads the minutes as the primary evidence that the Committee deliberated rather than ratified. An extract-only office would not know the minutes carry that weight. |
| 2.3 | absent | An unattainable objective is escalated to the Board as an amendment question | This is the operative instruction when the return objective cannot be met. The IPS forbids resolving it inside the portfolio by taking risk the Statement does not permit. The extract's 'best efforts' ranking alone does not say where the finding goes. |
| 2.5 | absent | Capital market assumptions from multiple named houses, with dispersion disclosed | The IPS states plainly that a single-source assumption is not acceptable and that the Committee is told which houses, of what vintage, and what the dispersion is. An extract-only office could table one forecast. |
| 3.3 | partial | A recommendation claiming both objectives are satisfied has not understood one of them | The extract asks which objective is being given ground on. The IPS goes further and rejects the answer 'neither'. |
| 3.4 | absent | The campaign inflow is staged into policy weights and is not timed against a market view | The extract records the USD 60m inflow but not the prohibition on timing it. Tactically staging that inflow would breach the IPS while looking compliant against the extract. |
| 3.5 | partial | Board exclusions are assessed at the vehicle level and incidental index exposure is disclosed rather than deemed compliant by silence | This is the most consequential omission. The extract lists the exclusions with no disclosure obligation, so a broad index vehicle carrying incidental tobacco or thermal coal exposure would pass silently. Every equity vehicle in the opportunity set is a broad index vehicle. |
| 3.5 | partial | Leverage is defined as gross exposure exceeding net asset value | The extract says 'no leverage at the fund level (UBTI)' without a testable definition. The IPS gives one, which is what makes it checkable by code. |
| 4.3 | absent | Presentation follows GIPS, including risk statistics for the benchmark and blended-benchmark disclosure | The extract carries no reporting standard at all. An extract-only office would not know it must show the benchmark's risk statistics and not only the portfolio's. |
| 4.5 | absent | Rebalancing on band breach rather than calendar, corridors set by volatility and cost, annual reset | The extract carries the 50bps minimum trade but not the rebalancing policy it interacts with, nor the requirement that the Committee be told what the corridors are and what determined them. |
| Assumption | What was assumed | Why, and what it costs |
|---|---|---|
| Study window | 1 July 2021 to 30 June 2026, sixty monthly observations, twenty quarterly meetings, fiscal years ending 30 June | The window is a parameter, not a constant. It is defined once in taa/config.py and every stage reads it from there, so changing it and rerunning reproduces the whole study on the new window with no other edit. Sixty observations is very thin: the standard error on a Sharpe ratio at this sample size is about ±0.45, which is wider than any result in this report. |
| Regimes in the window | One tightening cycle and one recovery, at most two regimes | A five-year window contains one or two regimes. Every statistic here is conditioned on a period containing the 2022 drawdown and the recovery that followed, and would look materially different on a window that excluded either. Shortening the window to three years turns the five-year active return of +22bps a year into (6)bps a year, which is the same programme measured differently. |
| Benchmark construction | Policy portfolio at IPS 4.1 weights, rebalanced monthly | Rebalancing frequency is a GIPS 24.C.27 disclosure item and this choice is not neutral. Tested after the fact: a never-rebalanced policy portfolio returns 8.56% a year over this window against 8.26% for the monthly-rebalanced blend, so the benchmark used here is 29.7bps a year EASIER to beat and the reported active return is correspondingly flattered. The report originally asserted the opposite without testing it. |
| Total return series | Yahoo Finance adjusted closes, dividends reinvested | Adjusted closes are restated retroactively when a dividend is paid, so the level of a past adjusted close is not what a screen showed that day. The return between two past dates is the correct total return, and no return in this study is computed across the as-of boundary, so no look-ahead is introduced. |
| Investable dates | Each line starts at its own vehicle's listing date | IPS 4.1 binds. Over this window every line was investable throughout, so nothing is spliced. On a longer window the point-in-time layer refuses a line before its vehicle existed rather than silently substituting index history. |
| Transaction costs | One-way, 1.5bps on US equity to 25bps on commodities | Issuer-published 30-day median bid-ask spreads under SEC Rule 6c-11, marked up for size and for premium-discount behaviour. Figures quoted for institutional single-stock trading do not apply and are not used. |
| Fees | Trading costs modelled; the 0.40% IPS 3.2 cost load not deducted from reported returns | The 0.40% is an office and custody cost inside the return requirement rather than a trading cost. GIPS 24.A.1.b requires a net-of-fees presentation for the composite; the returns here are net of trading and gross of that load, and the distinction is stated rather than assumed. |
| Reconciliation rule, historical record | Equal weight on the two desk allocations | Applied unchanged across all twenty meetings. Equal weight because neither desk has demonstrated skill that would justify preferring it. Every historical decision is mechanical and the report says so rather than inventing deliberation. |
This report is set in the Coldbrook Capital design system, the house system shipped in ds/ in this folder: ds/colors_and_type.css holds the tokens and ds/preview/ ships the component previews. Colours are the token values from that stylesheet rather than a remembered approximation, and the components follow their previews. Three departures are recorded. The stylesheet opens with a Google Fonts import, which is dropped so these documents render identically from disk with no network. The stylesheet sets --navy: #0C1E48 while every preview and the written brand document use #0A1B3D, and the stylesheet's own --up token is #0A1B3D described as "same as navy", which it no longer is; the tokens file governs and the drift is reported rather than silently resolved. And the wordmark is Ashcroft's, since Coldbrook is the fictional firm the system was authored for and putting its mark on another institution's document would be wrong. There is no red or green anywhere in these documents: direction is carried by navy against clay, by parentheses on negatives, and by position relative to a zero axis.
Every claim taken from the literature, with a source a reader can open, marked according to whether the desk verified it in session or recalled it. The distinction matters: a training cutoff makes a recalled number a hypothesis rather than a fact, and several of these have moved.
40 claims from the Systematic desk, 32 verified in session and 8 recalled (80% verified).
| Claim | Value | Source | Status |
|---|---|---|---|
| The classic equity-premium predictors have predicted poorly both in-sample and out-of-sample, are unstable, and would not have helped a real-time inve | negative OOS R2 for essentially the whole se | Welch & Goyal (2008), Review of Financial Studies 21(4), 1455-1508, abstract | VERIFIED |
| In Welch-Goyal Table 1, every one of the twelve classic monthly predictors has a negative full-sample out-of-sample R-squared against the expanding hi | OOS R2 from -1.78% (earnings-price) to -27.1 | Welch & Goyal (2008), RFS, Table 1, full-sample columns | VERIFIED |
| Re-examining 29 predictors from 26 papers published after 2008 plus the original 17, through end-2021, most do not survive. | more than one-third fail even in-sample; of | Goyal, Welch & Zafirov (2024), 'A Comprehensive 2022 Look at the Empirical Performance of Equity Premium Prediction', RFS 37(11), 3490-3557 | VERIFIED |
| Imposing two weak restrictions (sign restriction on the regression coefficient, non-negativity of the equity-premium forecast) makes most predictors b | small but positive OOS R2 for most restricte | Campbell & Thompson (2008), RFS 21(4), 1509-1531; NBER WP 11468 | VERIFIED |
| The correct yardstick for an out-of-sample R-squared is the squared Sharpe ratio at the same frequency, because the proportional increase in a mean-va | monthly squared Sharpe S2 = 1.2% (monthly SR | Campbell & Thompson (2008), Section 2, equations (8) and (9) | VERIFIED |
| The utility gain to a mean-variance investor from timing expected returns is about 35% of lifetime utility. | 35% | Campbell & Thompson (2008), as reported and benchmarked against by Moreira & Muir | VERIFIED |
| Large-scale replication of the published anomaly literature with microcaps controlled finds most anomalies fail. | of 452 anomalies, 65% fail |t|>1.96 (replica | Hou, Xue & Zhang (2020), 'Replicating Anomalies', RFS 33(5), 2019-2133 | VERIFIED |
| Multiple testing across the published factor zoo implies a much higher significance hurdle than t = 2.0, and a large fraction of published factors are | 316 factors catalogued from 313 papers; reco | Harvey, Liu & Zhu (2016), '...and the Cross-Section of Expected Returns', RFS 29(1), 5-68; NBER WP 20592 | VERIFIED |
| Published cross-sectional predictors decay out of sample and decay further after publication. | returns 26% lower out-of-sample (pre-publica | McLean & Pontiff (2016), 'Does Academic Research Destroy Stock Return Predictability?', Journal of Finance 71(1), 5-32 | VERIFIED |
| A Bayesian hierarchical model of factor replication that pools information across factors and countries reaches the opposite conclusion to Hou-Xue-Zha | 82.4% replication rate in the US and 82.4% g | Jensen, Kelly & Pedersen (2023), 'Is There a Replication Crisis in Finance?', Journal of Finance 78(5), 2465-2518 | VERIFIED |
| The gap between a 35% and an 82% replication rate is methodological and can be decomposed step by step, which is the actual finding. | 35% (HXZ) -> 55.6% (longer sample, 15 added | Jensen, Kelly & Pedersen (2023), Figure 1 and surrounding text | VERIFIED |
| Out-of-sample decay of published cross-sectional predictors is modest in the first years after the original sample ends. | 74% of in-sample long-short return persists | Chen & Zimmermann (2022), 'Publication Bias in Asset Pricing Research' | VERIFIED |
| The standard reconciliation of Hou-Xue-Zhang with other meta-studies (microcap de-emphasis) is wrong; the real driver is misclassification of failures | only about 26% of Hou-Xue-Zhang's long-short | Chen & Zimmermann (2022), 'Publication Bias in Asset Pricing Research', Section 2.1 | VERIFIED |
| Open Source Asset Pricing reproduces published cross-sectional predictors at near-perfect rates in-sample. | 98% of 161 clearly-significant original char | Chen & Zimmermann, 'Open Source Cross-Sectional Asset Pricing', Critical Finance Review | RECALLED |
| Time-series momentum: the past 12-month excess return positively predicts the next month's return across futures markets. | 58 liquid futures and forward instruments; a | Moskowitz, Ooi & Pedersen (2012), Journal of Financial Economics 104(2), 228-250 | VERIFIED |
| The statistical evidence for time-series momentum does not survive proper inference. | asset-by-asset regressions show little TSM i | Huang, Li, Wang & Zhou (2020), 'Time series momentum: Is it there?', JFE 135(3), 774-794 | VERIFIED |
| Carry predicts returns cross-sectionally and in the time series across asset classes, and is not subsumed by known predictors. | average carry-strategy Sharpe 0.74 across ni | Koijen, Moskowitz, Pedersen & Vrugt (2018), 'Carry', JFE 127(2), 197-225; NBER WP 19325 | VERIFIED |
| Value and momentum premia appear consistently across eight markets and asset classes with a strong common factor structure, and are negatively correla | eight markets and asset classes; three-facto | Asness, Moskowitz & Pedersen (2013), 'Value and Momentum Everywhere', Journal of Finance 68(3), 929-985 | VERIFIED |
| When returns are regressed on a lagged persistent stochastic regressor such as dividend yield, the OLS estimator is biased in finite samples because t | over 1927-96 the bias equals roughly one-thi | Stambaugh (1999), 'Predictive Regressions', JFE 54(3), 375-421 | RECALLED |
| Long-horizon predictability evidence built from overlapping regressions is not independent evidence, because the estimators are almost perfectly corre | analytical correlation of 99% between 1- and | Boudoukh, Richardson & Whitelaw (2008), 'The Myth of Long-Horizon Predictability', RFS 21(4), 1577-1605 | VERIFIED |
| Scaling factor exposure by the inverse of last month's realised variance produces positive alphas and higher Sharpe ratios across many factors. | market strategy: annualised alpha 4.86%, app | Moreira & Muir (2017), 'Volatility-Managed Portfolios', Journal of Finance 72(4), 1611-1644; NBER WP 22208 | VERIFIED |
| Volatility management does not survive real-time implementation. | across 103 equity strategies, reasonable out | Cederburg, O'Doherty, Wang & Yan (2020), 'On the performance of volatility-managed portfolios', JFE 138(1), 95-117 | VERIFIED |
| Momentum's risk is time-varying and forecastable, and targeting constant volatility on the momentum factor removes the crashes. | targeting 12% constant volatility virtually | Barroso & Santa-Clara (2015), 'Momentum has its moments', JFE 116(1), 111-120 | VERIFIED |
| Volatility targeting helps risk assets and does little for anything else, and its main benefit is tail-shape rather than Sharpe. | 60 assets, daily data from 1926 to 2017; Sha | Harvey, Hoyle, Korgaonkar, Rattray, Sargaison & Van Hemert (2018), 'The Impact of Volatility Targeting', Journal of Portfolio Management 45(1), 14-33 | RECALLED |
| The standard volatility-timing implementation contains a look-ahead bias, and correcting it destroys the result. | after correcting the bias the maximum drawdo | Liu, Tang & Zhou (2019/2020), 'Volatility-Managed Portfolio: Does It Really Work?', Journal of Portfolio Management 46(1), 38-51 | RECALLED |
| Conditional variance is autoregressive: volatility clusters and is forecastable from its own history. | ARCH (Engle 1982, Econometrica) and GARCH (B | Engle (1982), Econometrica 50(4), 987-1007; Bollerslev (1986), Journal of Econometrics 31(3), 307-327 | RECALLED |
| A simple heterogeneous autoregressive model of realised volatility forecasts volatility with high accuracy out of sample. | out-of-sample Mincer-Zarnowitz R2 for S&P 50 | Corsi (2009), 'A Simple Approximate Long-Memory Model of Realized Volatility', Journal of Financial Econometrics 7(2), 174-196; author's lecture slides, SNS Pisa 2010 | VERIFIED |
| Volatility and returns are not comparably forecastable, and the gap is roughly two orders of magnitude. | R2_oos of roughly 0.76 to 0.82 for realised | Derived from Corsi (2009) and Welch & Goyal (2008) / Campbell & Thompson (2008) | VERIFIED |
| The information ratio of an unconstrained active programme is approximately the information coefficient times the square root of breadth. | IR = IC x sqrt(BR) | Grinold (1989), 'The Fundamental Law of Active Management', Journal of Portfolio Management 15(3), 30-37 | RECALLED |
| Portfolio constraints break the fundamental law; the correction is a transfer coefficient, the cross-sectional correlation between risk-adjusted activ | IR = TC x IC x sqrt(BR); their simulation st | Clarke, de Silva & Thorley (2002), Financial Analysts Journal 58(5), 48-66, as restated and quoted in Zhou, 'The Fundamental Law of Active Management: Redux' | VERIFIED |
| Measured transfer coefficients for a realistically constrained long-only portfolio are well below one, and the long-only constraint is the single larg | TC = 0.332 with all constraints imposed; 0.6 | Clarke, de Silva & Sapra (2004), 'Toward More Information-Efficient Portfolios', Journal of Portfolio Management, Fall 2004, 54-63 | VERIFIED |
| Breadth must be adjusted downward for correlation between bets. | BR_eff = N / (1 + rho x (N - 1)) | Buckle (2004), 'How to calculate breadth: An evolution of the fundamental law of active portfolio management', Journal of Asset Management 4(6), 393-405 | RECALLED |
| The effective-breadth haircut is contested: modelling the full portfolio-construction problem suggests correlation between returns and between alphas | return correlations and alpha correlations b | Sneddon (2020), 'Strategy Design and the Fallacies of Breadth', Journal of Asset Management 21, 626-635; Northfield webinar, March 2021 | VERIFIED |
| Benchmark scale for information coefficients. | a good forecaster has IC = 0.05, a great for | Grinold & Kahn, 'Active Portfolio Management', chapter 10 (Forecasting), as reproduced in Yan's chapter-by-chapter study notes | VERIFIED |
| Grinold and Kahn themselves state that quarterly benchmark timing has too little breadth to produce a respectable information ratio. | an independent benchmark-timing forecast eve | Grinold & Kahn, 'Active Portfolio Management', chapter 15 (Benchmark Timing), as reproduced in Yan's study notes | VERIFIED |
| The fundamental law understates true active risk because it ignores variation in the information coefficient itself; realised active risk is often mat | IR is better estimated as mean(IC) / std(IC) | Qian & Hua, 'Active Risk and Information Ratio', Journal of Investment Management | VERIFIED |
| Standard error of an estimated Sharpe ratio under IID returns. | SE(SR) = sqrt((1 + SR^2 / 2) / T); at T = 60 | Lo (2002), 'The Statistics of Sharpe Ratios', Financial Analysts Journal 58(4), 36-52, equation (9) and Table 1 | VERIFIED |
| Definition of the out-of-sample R-squared used throughout this literature. | R2_oos = 1 - sum(r_t - rhat_t)^2 / sum(r_t - | Campbell & Thompson (2008), Section 1; identical benchmark to Welch & Goyal (2008) | VERIFIED |
| Nested-model forecast comparison requires adjusting the MSPE difference for the noise the larger model introduces by estimating parameters that are ze | fhat_{t+J} = (y_{t+J} - yhat1)^2 - [(y_{t+J} | Clark & West (2007), 'Approximately normal tests for equal predictive accuracy in nested models', Journal of Econometrics 138(1), 291-311, equations (2.1) and (3.15) | VERIFIED |
| Blended all-in one-way execution cost for the nine implementation vehicles. | approximately 8bps one-way, weighted across | Desk assumption, calibrated to published ETF bid-ask spread ranges | RECALLED |
| Line | Policy | Quantitative desk | Active | Macro desk | Active |
|---|---|---|---|---|---|
| % | % | bps | % | bps | |
| US equity | 38 | 32.46 | (554) | 36.00 | (200) |
| Developed ex-US | 20 | 21.80 | 180 | 20.00 | 0 |
| Emerging markets | 12 | 14.44 | 244 | 11.00 | (100) |
| US Treasury duration | 12 | 22.00 | 1,000 | 5.50 | (650) |
| US investment grade | 8 | 3.00 | (500) | 7.00 | (100) |
| US high yield | 5 | 0.00 | (500) | 3.50 | (150) |
| Commodities | 3 | 3.11 | 11 | 8.00 | 500 |
| Listed real estate | 2 | 3.19 | 119 | 2.00 | 0 |
| T-bills | 0 | 0.00 | 0 | 7.00 | 700 |
| Total | 100 | 100.00 | 100.00 |
Quantitative desk rationale, as submitted: The desk recommends the policy portfolio. The composite signal has a pooled out-of-sample R2 of -1.61 per cent against the expanding historical mean, negative on 6 of 8 lines, with a Clark-West t of -0.77. Across the five signals and eight risky lines, 9 of 40 cells are positive. On that evidence the estimator does not earn active risk, and the size of a tilt should follow the evidence for it rather than the size of the budget available to express it. The constrained allocation reported here is the mechanical output of the optimiser run at the full budget; it is the model's answer and it is wh
Macro desk rationale, as submitted: Overheat. Growth is expansion on the 30 June vintage (growth score 4 of a possible 4: real GDP 2.09% annualised for 2026 Q1, payrolls averaging 188k over three months, a Sahm gap of 0.10pp, industrial production 1.67% year on year), and inflation is above target and firming (core PCE 3.41% year on year with a 3.52% three-month annualised run rate). The position that follows is not a growth call. It is a call on the price of two things. First, policy: the 3-month bill at 3.87% against core PCE at 3.41% is a real policy rate of 0.46%, below any plausible neutral, so the stance is neutral rather
The Capital Markets desk fetched every house forecast in session and each carries a URL. Two inputs are marked recalled and stated as such at the point of use: the long-run real earnings growth assumption in the bottom-up equity cross-check, and parts of the historical context around prior CAPE peaks. The Systematic desk fetched the primary papers rather than summaries of them where the paper was reachable, and marks eight of forty claims as recalled. Where a desk could not reach a primary source it says which secondary source it used instead, which happened once, for Research Affiliates, whose own tool requires a login and whose figures were taken from Morningstar.
No paywalled source was used anywhere in this study and no credential was required. The one place a subscription would have improved the work is a clean forward earnings yield for the equity risk premium proxy, which the Macro desk records as out of scope by design and substitutes a dividend-yield construction for.
The five-year decision record, twenty quarterly entries with reasons, binding constraints and compliance results, is a separate document: decision_record.html. The interactive record is dashboard.html, which opens from disk with no network. The methods notebook, tying every method to the paper it comes from, is methods.ipynb. Every algorithm is a runnable Python file in taa/, the desk papers are in desks/, and the verification artifacts are in tests/.
Two supplementary notes sit in outputs/. Three things the office did that were worth watching records how six concurrent desks behaved as a system, with the clock time of each episode. The three decisions that earned the most is an outcome-selected analysis of the best three of the twenty, written with hindsight deliberately and labelled as such, and it does not change the recommendation.
Before any of this is published or quoted outside the Committee, read AUDIT.md. It records what in this study can and cannot be trusted: that the five-year record is a simulation of a rule rather than a track record, that the performance numbers cannot be separated from zero and change sign with the measurement window, and that the committee minutes are a construct whose reasoning is genuine and whose meeting is not.