Models are now a Composite of drift, volatility and (optional) shape components, so an experiment can swap one part against a fixed control. btcmodel/experiments.py holds seven experiments built from the ideas in the old branches (catalogued in docs/2024-ideas.md), each with its hypothesis and source, and a verdict rule fixed before anything ran. `just ab` runs them on development data. Results: - Shrinking the cycle drift, and a power-law trend (plain or reverting), beat their controls. The power law beats the random walk by 53-63% at 3-4 years with unbiased outcomes, so it is promoted to MODELS. - Every alternative volatility estimate (EWMA blends, other windows, reversion to a level or trend) is worse than the trailing 365-day window. Cycle-dependent volatility, heavy tails and stretched cycle phase show no reliable effect.
141 lines
6.4 KiB
Markdown
141 lines
6.4 KiB
Markdown
# Bitcoin Price Model
|
||
|
||
<p align="center">
|
||
<img src="https://img.izismile.com/img/img5/20120417/640/i_have_no_idea_what_im_doing_meme_640_07.jpg" alt="I have no idea what I'm doing" />
|
||
</p>
|
||
|
||
**Don't take this seriously. It's all in good fun.**
|
||
|
||
## The 2024 edition
|
||
|
||
In November 2024 I asked Claude (3.6 Sonnet, via copy and paste in Claude Web)
|
||
to help build a Bitcoin price model. After a lot of branching it produced ~2000
|
||
lines of cycle analysis, "market fundamentals", era adjustments and Monte Carlo,
|
||
and forecasts like this one (as of 2024-11-14):
|
||
|
||

|
||
|
||
In September 2026 a newer Claude scored the forecasts it made on 2024-11-27
|
||
against what actually happened:
|
||
|
||
| | 2024 forecast (median) | Actual |
|
||
|---|---|---|
|
||
| Cycle top | 2025-10-21, $199K | 2025-10-06, $124.7K |
|
||
| 2026-06-30 | $133K (95% floor $65K) | $58.5K |
|
||
| 2026-09-23 | $124K | $84.4K |
|
||
|
||
It called the timing of the top within 15 days, 19 months ahead, but put the
|
||
level ~60% too high. "The price stays at $92.7K" was a better forecast (18%
|
||
mean error against 68%). Its "95%" intervals were really ~81% intervals past a
|
||
year out, thanks to a fudge factor that narrowed them at long horizons.
|
||
|
||
That code lives in the jj/git history. This is the rewrite.
|
||
|
||
## How it works now
|
||
|
||
A forecast is a probability distribution of the price at each horizon, not a
|
||
line. Every model produces quantiles of log price, and every model is scored
|
||
the same way:
|
||
|
||
- **Walk-forward.** Every 30 days from 2014 on, each model sees only the data up
|
||
to that day and forecasts 1 month to 4 years ahead.
|
||
- **CRPS.** Each forecast is scored against what happened with the continuous
|
||
ranked probability score, in log-price units (0.1 ≈ "typically 10% off"). It
|
||
rewards being sharp and being calibrated at once, and can't be gamed by
|
||
narrowing or widening intervals.
|
||
- **Baselines.** Skill is reported relative to a zero-drift random walk, with a
|
||
90% block-bootstrap interval. Nearby forecasts overlap heavily, so the report
|
||
also shows `windows`: the number of genuinely independent outcomes.
|
||
- **Holdout.** Development only sees data up to 2024-11-26, the last day the
|
||
2024 model saw. Everything after it is held out, and is scored only by
|
||
`just holdout`, once per round of model changes. (Caveat: we already know
|
||
roughly what happened in 2025-26, so it isn't perfectly blind.)
|
||
|
||
### Models
|
||
|
||
- `random_walk`: zero drift; "it stays about here, give or take".
|
||
- `drift_rw`: drift equal to the last four years' average; "it keeps doing what
|
||
it did last cycle".
|
||
- `cycle`: the 2024 model's one real idea. Expected return depends on the day of
|
||
the halving cycle, estimated from past cycles, with recent cycles weighted
|
||
more.
|
||
- `powerlaw`: log price grows linearly in log time since genesis, so growth
|
||
keeps slowing. Fitted walk-forward; the exponent has stayed between 5.4 and
|
||
6.0 in every fit since 2014.
|
||
|
||
### Findings so far (development data)
|
||
|
||
- `cycle` loses to the random walk at every horizon, and so does every setting
|
||
tried (recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days).
|
||
The level is the problem: each cycle has grown less than the last (log
|
||
return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any average of past
|
||
cycles overshoots.
|
||
- `powerlaw` models exactly that, and it is the first model to beat the random
|
||
walk with some confidence: +53% and +63% skill at 3 and 4 years, with
|
||
unbiased outcomes (mean PIT 0.51). Only 3-4 independent windows back that
|
||
up, the functional form is famous *because* it fits Bitcoin's history, and
|
||
the holdout hasn't been run yet.
|
||
- Its intervals are too wide at long horizons (the 80% interval held every
|
||
3-year outcome), because it treats deviations from the trend as permanent.
|
||
|
||
### A/B tests of the 2024 ideas
|
||
|
||
`just ab` runs the experiments in `btcmodel/experiments.py`: ideas salvaged
|
||
from the old branches (catalogued in [docs/2024-ideas.md](docs/2024-ideas.md)),
|
||
each a control plus variants that change one component. Hypotheses and the
|
||
verdict rule were written down before anything ran. With ~100 comparisons,
|
||
expect a few flukes either way.
|
||
|
||
| Experiment | Idea | Verdict |
|
||
|---|---|---|
|
||
| shrink-cycle | scale the cycle drift by 0.25/0.5/0.75 | better, all three: it fixes the level crudely |
|
||
| diminishing-returns | power-law trend, optionally reverting to it, or the cycle shape rescaled to it | better (both power-law variants); cycle shape on the power law inconclusive |
|
||
| vol-window | EWMA blends, shorter or longer windows | worse: the plain 365-day window wins |
|
||
| vol-reversion | volatility reverting to a long-run level or falling trend | worse |
|
||
| cycle-vol | volatility by cycle position | inconclusive (no effect) |
|
||
| tails | Student-t, or empirical horizon-level shape | Student-t worse; empirical +5% at 1 month only |
|
||
| cycle-phase | align cycles by fraction elapsed, not days | inconclusive |
|
||
|
||
Next: a power law whose deviations revert with bounded variance, and a
|
||
registered test of whether the cycle's timing adds anything on top of it.
|
||
|
||
## Usage
|
||
|
||
With Nix:
|
||
|
||
```sh
|
||
nix develop
|
||
just update # fetch new daily prices from Coinbase
|
||
just backtest # score models on development data -> output/backtest/
|
||
just forecast # forecast from the latest price -> output/forecast/
|
||
just ab # run A/B experiments -> output/ab/
|
||
just test
|
||
just holdout # score on held-out outcomes; sparingly
|
||
```
|
||
|
||
Without Nix, any Python 3.13 with numpy, pandas 3, scipy and matplotlib works:
|
||
`python -m btcmodel --help`.
|
||
|
||
## Layout
|
||
|
||
```
|
||
btcmodel/
|
||
data.py price loading (Investing.com archive + Coinbase), dev cutoff
|
||
halving.py halving calendar, position in cycle
|
||
forecast.py Forecast (quantiles of log price), CRPS, PIT
|
||
evaluate.py walk-forward backtest and summary
|
||
experiments.py A/B tests: hypothesis, control, variants, verdict rule
|
||
models/ drift, volatility and shape components; register models in __init__.py
|
||
plots.py fan chart, skill and calibration charts
|
||
```
|
||
|
||
A model is any object with a `name` and
|
||
`forecast(history, horizons) -> Forecast`. `history` is a date-indexed frame
|
||
holding everything up to the forecast origin and nothing after it. Other data
|
||
sources (hash rate, on-chain metrics, macro series) can join it as extra
|
||
columns. Anything published with a lag must be shifted to the date it was
|
||
actually available, or the backtest will quietly cheat.
|
||
|
||
Price data: [Investing.com](https://www.investing.com/crypto/bitcoin/historical-data)
|
||
through 2024-11-26, then Coinbase Exchange daily closes (UTC).
|