HISTORICAL EVALUATION
What the historical tests establish
Claims show useful warning information; combined forecasts remain unvalidated
Weekly claims added useful recession-onset information in this retrospective test. Simply blending claims, EBP and OFR did not improve discrimination over claims alone.
Monthly recession-onset diagnostics using current-revised claims, EBP and OFR histories with assumed publication delays. The complete ECO I synthesis was not evaluated. A separately frozen follow-on evaluates exact live claims-change formulas using the same already-examined holdout.
These are retrospective tests using revised data and assumed publication delays. They do not validate ECO I’s complete assessment, predict a bubble burst, or estimate a probability.
Claims and credit: longer history
Training: 31 Jan 1975 to 31 Dec 1999. 268 eligible expansion months tested. Each warning asks whether recession onset follows within 12 months.
| Candidate rule | Discrimination (AUROC) | Onsets warned | Onsets missed | False warning episodes |
|---|---|---|---|---|
| Claims | 0.755 | 3 of 3 | 0 | 7 |
| Ebp | 0.624 | 1 of 3 | 2 | 9 |
| Claims and EBP | 0.708 | 2 of 3 | 1 | 11 |
| Never warning | 0.500 | 0 of 3 | 3 | 0 |
| Always warning | 0.500 | 3 of 3 | 0 | 1 |
AUROC measures discrimination between pre-onset and other expansion months. It is neither a crash probability nor the percentage of forecasts that were correct.
Claims, credit and OFR: common history
Training: 31 Jan 2000 to 31 Dec 2006. 192 eligible expansion months tested. Each warning asks whether recession onset follows within 12 months.
| Candidate rule | Discrimination (AUROC) | Onsets warned | Onsets missed | False warning episodes |
|---|---|---|---|---|
| Claims | 0.751 | 2 of 2 | 0 | 6 |
| Ebp | 0.413 | 0 of 2 | 2 | 3 |
| Ofr | 0.429 | 2 of 2 | 0 | 15 |
| Claims and EBP | 0.666 | 0 of 2 | 2 | 4 |
| Claims EBP OFR | 0.594 | 1 of 2 | 1 | 4 |
| Never warning | 0.500 | 0 of 2 | 2 | 0 |
| Always warning | 0.500 | 2 of 2 | 0 | 1 |
AUROC measures discrimination between pre-onset and other expansion months. It is neither a crash probability nor the percentage of forecasts that were correct.
Follow-on: displayed claims changes
Training: 31 Jan 1975 to 31 Dec 1999. 268 eligible expansion months tested. Each warning asks whether recession onset follows within 12 months.
| Candidate rule | Discrimination (AUROC) | Onsets warned | Onsets missed | False warning episodes |
|---|---|---|---|---|
| Live claims changes | 0.776 | 3 of 3 | 0 | 13 |
| Original claims pressure | 0.755 | 3 of 3 | 0 | 7 |
| Never warning | 0.500 | 0 of 3 | 3 | 0 |
| Always warning | 0.500 | 3 of 3 | 0 | 1 |
AUROC measures discrimination between pre-onset and other expansion months. It is neither a crash probability nor the percentage of forecasts that were correct.
How this affects the application
Use claims in labor context, EBP in corporate-credit context and OFR in stress explanation. Do not promote the naive blend or increase predictive confidence from this diagnostic.
- With training fixed to 1975–1999, claims warned within 12 months before all 3 recession onsets in the 2000–August 2024 test, with 7 false warning episodes and AUROC 0.755. A warning before an event alone does not establish a reliable forecast.
- On that same long cohort, EBP alone had AUROC 0.624 and claims plus EBP 0.708, below claims alone. EBP alone warned before 1 of 3 onsets; the combination warned before 2 of 3.
- On the common 2007–August 2024 cohort, claims had AUROC 0.751, claims plus EBP 0.666, and claims plus EBP plus OFR 0.594. Adding these financial inputs to this simple rule did not improve ordering of pre-onset months.
- The follow-on exact live claims-change rule had AUROC 0.776, warnings before all 3 onsets, and 13 false warning episodes among 18 warning episodes. The original pressure rule had AUROC 0.755 and 7 false episodes on the identical cohort. This is not independent confirmation or demonstrated superiority; the same holdout had already been inspected.
- Only 3 recession onsets support the long test and 2 support the OFR test. Overlapping monthly outcomes are not independent events; warnings before 2020 do not establish that the pandemic was anticipated.
Capture authentic historical release vintages and compare target-specific rules with additional simple benchmarks on frozen common cohorts before changing forecast claims.
Download all inputs, results and the frozen specification (JSON) →Command line: eco-i evaluate. Reproduce a study with eco-i evaluate --input study.json; the JSON report contains the reproduction specification and input observations.