Round 2: reverting power law and cycle-on-powerlaw tests; one-time holdout run.
Add TrendReversionVol: deviations from the power-law trend follow a daily AR(1), so uncertainty levels off, optionally plus trend-parameter uncertainty with an autocorrelation-adjusted effective sample size. Two experiments, run under the unchanged verdict rule: - powerlaw-ou: +21% to +45% vs powerlaw at 2-4 years, but slightly negative point estimates at 1 month make it inconclusive. - cycle-on-powerlaw: inconclusive (+18% at 2 years, negative elsewhere). The holdout (outcomes after 2024-11-26) was scored once, for the four candidates fixed beforehand. powerlaw is the best long-horizon forecast (+45% and +58% vs the random walk at 2 and 3 years); nothing beats the random walk inside a year; cycle fails badly. Results are in the README.
This commit is contained in:
@@ -96,8 +96,32 @@ expect a few flukes either way.
|
||||
| tails | Student-t, or empirical horizon-level shape | Student-t worse; empirical +5% at 1 month only |
|
||||
| cycle-phase | align cycles by fraction elapsed, not days | inconclusive |
|
||||
|
||||
Next: a power law whose deviations revert with bounded variance, and a
|
||||
registered test of whether the cycle's timing adds anything on top of it.
|
||||
Round 2 tested the power law against two refinements (the verdict rule was
|
||||
unchanged, and the holdout candidates were fixed before it ran):
|
||||
|
||||
| Experiment | Idea | Verdict |
|
||||
|---|---|---|
|
||||
| powerlaw-ou | deviations from the trend revert, so uncertainty levels off; optionally plus trend-parameter uncertainty | inconclusive: +21% to +45% at 2-4 years (intervals above zero), but -0.3% and -1% at 1 month fail the "never negative" clause. Bands too narrow without parameter uncertainty, too wide with it |
|
||||
| cycle-on-powerlaw | the cycle's timing, rescaled to the power-law level | inconclusive: +18% at 2 years, slightly negative at short horizons and 4 years |
|
||||
|
||||
### Holdout (run once, 2026-09-24)
|
||||
|
||||
Scored on outcomes after 2024-11-26 for the four candidates fixed in advance
|
||||
(`random_walk`, `drift_rw`, `cycle`, `powerlaw`). The holdout is ~22 months
|
||||
long, so at 2-4 years it is essentially one outcome period seen from several
|
||||
origins (1.4-1.9 windows), and the bootstrap intervals there mean little.
|
||||
|
||||
| Skill vs random walk | 1mo | 3mo | 6mo | 1y | 2y | 3y | 4y |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| drift_rw | -0% | +1% | +5% | +14% | +31% | +53% | -65% |
|
||||
| cycle | -14% | -8% | -25% | -104% | -46% | -21% | -179% |
|
||||
| powerlaw | -1% | -7% | -5% | +3% | +45% | +58% | +15% |
|
||||
|
||||
Consistent with development: nothing beats the random walk inside a year,
|
||||
`cycle` fails badly, and `powerlaw` is the best long-horizon forecast, unbiased
|
||||
at 2-3 years (mean PIT 0.53-0.56) but with intervals too wide (its 80% interval
|
||||
held every 2- and 3-year outcome). The holdout is now spent for these models;
|
||||
a new model needs new data to be tested honestly.
|
||||
|
||||
## Usage
|
||||
|
||||
|
||||
Reference in New Issue
Block a user