Composable models and A/B tests of the 2024 ideas; add powerlaw.
Models are now a Composite of drift, volatility and (optional) shape components, so an experiment can swap one part against a fixed control. btcmodel/experiments.py holds seven experiments built from the ideas in the old branches (catalogued in docs/2024-ideas.md), each with its hypothesis and source, and a verdict rule fixed before anything ran. `just ab` runs them on development data. Results: - Shrinking the cycle drift, and a power-law trend (plain or reverting), beat their controls. The power law beats the random walk by 53-63% at 3-4 years with unbiased outcomes, so it is promoted to MODELS. - Every alternative volatility estimate (EWMA blends, other windows, reversion to a level or trend) is worse than the trailing 365-day window. Cycle-dependent volatility, heavy tails and stretched cycle phase show no reliable effect.
This commit is contained in:
@@ -59,20 +59,45 @@ the same way:
|
||||
- `cycle`: the 2024 model's one real idea. Expected return depends on the day of
|
||||
the halving cycle, estimated from past cycles, with recent cycles weighted
|
||||
more.
|
||||
- `powerlaw`: log price grows linearly in log time since genesis, so growth
|
||||
keeps slowing. Fitted walk-forward; the exponent has stayed between 5.4 and
|
||||
6.0 in every fit since 2014.
|
||||
|
||||
### Findings so far (development data)
|
||||
|
||||
- Nothing beats the random walk with any confidence at any horizon. At 2-4
|
||||
years there are only 3-5 independent outcomes in the whole history.
|
||||
- `drift_rw` leads at 2-4 years (+13-18% skill, but the intervals span zero).
|
||||
- `cycle` loses to both at every horizon, and so does every setting tried
|
||||
(recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days). The cycle
|
||||
*shape* costs accuracy. The level is the problem: each cycle has grown less
|
||||
than the last (log return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any
|
||||
average of past cycles overshoots.
|
||||
- `cycle` loses to the random walk at every horizon, and so does every setting
|
||||
tried (recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days).
|
||||
The level is the problem: each cycle has grown less than the last (log
|
||||
return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any average of past
|
||||
cycles overshoots.
|
||||
- `powerlaw` models exactly that, and it is the first model to beat the random
|
||||
walk with some confidence: +53% and +63% skill at 3 and 4 years, with
|
||||
unbiased outcomes (mean PIT 0.51). Only 3-4 independent windows back that
|
||||
up, the functional form is famous *because* it fits Bitcoin's history, and
|
||||
the holdout hasn't been run yet.
|
||||
- Its intervals are too wide at long horizons (the 80% interval held every
|
||||
3-year outcome), because it treats deviations from the trend as permanent.
|
||||
|
||||
That points at diminishing returns as the structure worth modelling, e.g. a
|
||||
power-law trend, which is next.
|
||||
### A/B tests of the 2024 ideas
|
||||
|
||||
`just ab` runs the experiments in `btcmodel/experiments.py`: ideas salvaged
|
||||
from the old branches (catalogued in [docs/2024-ideas.md](docs/2024-ideas.md)),
|
||||
each a control plus variants that change one component. Hypotheses and the
|
||||
verdict rule were written down before anything ran. With ~100 comparisons,
|
||||
expect a few flukes either way.
|
||||
|
||||
| Experiment | Idea | Verdict |
|
||||
|---|---|---|
|
||||
| shrink-cycle | scale the cycle drift by 0.25/0.5/0.75 | better, all three: it fixes the level crudely |
|
||||
| diminishing-returns | power-law trend, optionally reverting to it, or the cycle shape rescaled to it | better (both power-law variants); cycle shape on the power law inconclusive |
|
||||
| vol-window | EWMA blends, shorter or longer windows | worse: the plain 365-day window wins |
|
||||
| vol-reversion | volatility reverting to a long-run level or falling trend | worse |
|
||||
| cycle-vol | volatility by cycle position | inconclusive (no effect) |
|
||||
| tails | Student-t, or empirical horizon-level shape | Student-t worse; empirical +5% at 1 month only |
|
||||
| cycle-phase | align cycles by fraction elapsed, not days | inconclusive |
|
||||
|
||||
Next: a power law whose deviations revert with bounded variance, and a
|
||||
registered test of whether the cycle's timing adds anything on top of it.
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -83,6 +108,7 @@ nix develop
|
||||
just update # fetch new daily prices from Coinbase
|
||||
just backtest # score models on development data -> output/backtest/
|
||||
just forecast # forecast from the latest price -> output/forecast/
|
||||
just ab # run A/B experiments -> output/ab/
|
||||
just test
|
||||
just holdout # score on held-out outcomes; sparingly
|
||||
```
|
||||
@@ -98,7 +124,8 @@ btcmodel/
|
||||
halving.py halving calendar, position in cycle
|
||||
forecast.py Forecast (quantiles of log price), CRPS, PIT
|
||||
evaluate.py walk-forward backtest and summary
|
||||
models/ one file per model family; register new ones in __init__.py
|
||||
experiments.py A/B tests: hypothesis, control, variants, verdict rule
|
||||
models/ drift, volatility and shape components; register models in __init__.py
|
||||
plots.py fan chart, skill and calibration charts
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user