Files
bitcoin-model/README.md
T
sam 67d016fe25 Forward test: record forecasts before their outcomes exist.
`snapshot` writes each tracked model's forecast quantiles at the seven
backtest horizons from the latest price to data/forecasts/<origin>.csv. It
refuses stale data (older than two days) and duplicate dates, so snapshots
can't be reconstructed after the fact; committing them dates them.
`forward` scores every recorded forecast whose target date has passed,
reusing the backtest's scoring (now factored out as evaluate.score).

Tracked: random_walk, drift_rw, cycle, powerlaw, plus powerlaw_ou,
powerlaw_ou_param and cycle_on_powerlaw, which development data couldn't
settle. `just weekly` runs update, snapshot and forward.

First snapshot: 2026-09-23 (BTC $84.4K). The first outcomes are due
2026-10-23.
2026-09-24 03:08:30 -07:00

184 lines
8.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Bitcoin Price Model
<p align="center">
<img src="https://img.izismile.com/img/img5/20120417/640/i_have_no_idea_what_im_doing_meme_640_07.jpg" alt="I have no idea what I'm doing" />
</p>
**Don't take this seriously. It's all in good fun.**
## The 2024 edition
In November 2024 I asked Claude (3.6 Sonnet, via copy and paste in Claude Web)
to help build a Bitcoin price model. After a lot of branching it produced ~2000
lines of cycle analysis, "market fundamentals", era adjustments and Monte Carlo,
and forecasts like this one (as of 2024-11-14):
![moooooon](docs/2024-forecast.png)
In September 2026 a newer Claude scored the forecasts it made on 2024-11-27
against what actually happened:
| | 2024 forecast (median) | Actual |
|---|---|---|
| Cycle top | 2025-10-21, $199K | 2025-10-06, $124.7K |
| 2026-06-30 | $133K (95% floor $65K) | $58.5K |
| 2026-09-23 | $124K | $84.4K |
It called the timing of the top within 15 days, 19 months ahead, but put the
level ~60% too high. "The price stays at $92.7K" was a better forecast (18%
mean error against 68%). Its "95%" intervals were really ~81% intervals past a
year out, thanks to a fudge factor that narrowed them at long horizons.
That code lives in the jj/git history. This is the rewrite.
## How it works now
A forecast is a probability distribution of the price at each horizon, not a
line. Every model produces quantiles of log price, and every model is scored
the same way:
- **Walk-forward.** Every 30 days from 2014 on, each model sees only the data up
to that day and forecasts 1 month to 4 years ahead.
- **CRPS.** Each forecast is scored against what happened with the continuous
ranked probability score, in log-price units (0.1 ≈ "typically 10% off"). It
rewards being sharp and being calibrated at once, and can't be gamed by
narrowing or widening intervals.
- **Baselines.** Skill is reported relative to a zero-drift random walk, with a
90% block-bootstrap interval. Nearby forecasts overlap heavily, so the report
also shows `windows`: the number of genuinely independent outcomes.
- **Holdout.** Development only sees data up to 2024-11-26, the last day the
2024 model saw. Everything after it is held out, and is scored only by
`just holdout`, once per round of model changes. (Caveat: we already know
roughly what happened in 2025-26, so it isn't perfectly blind.)
### Models
- `random_walk`: zero drift; "it stays about here, give or take".
- `drift_rw`: drift equal to the last four years' average; "it keeps doing what
it did last cycle".
- `cycle`: the 2024 model's one real idea. Expected return depends on the day of
the halving cycle, estimated from past cycles, with recent cycles weighted
more.
- `powerlaw`: log price grows linearly in log time since genesis, so growth
keeps slowing. Fitted walk-forward; the exponent has stayed between 5.4 and
6.0 in every fit since 2014.
### Findings so far (development data)
- `cycle` loses to the random walk at every horizon, and so does every setting
tried (recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days).
The level is the problem: each cycle has grown less than the last (log
return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any average of past
cycles overshoots.
- `powerlaw` models exactly that, and it is the first model to beat the random
walk with some confidence: +53% and +63% skill at 3 and 4 years, with
unbiased outcomes (mean PIT 0.51). Only 3-4 independent windows back that
up, the functional form is famous *because* it fits Bitcoin's history, and
the holdout hasn't been run yet.
- Its intervals are too wide at long horizons (the 80% interval held every
3-year outcome), because it treats deviations from the trend as permanent.
### A/B tests of the 2024 ideas
`just ab` runs the experiments in `btcmodel/experiments.py`: ideas salvaged
from the old branches (catalogued in [docs/2024-ideas.md](docs/2024-ideas.md)),
each a control plus variants that change one component. Hypotheses and the
verdict rule were written down before anything ran. With ~100 comparisons,
expect a few flukes either way.
| Experiment | Idea | Verdict |
|---|---|---|
| shrink-cycle | scale the cycle drift by 0.25/0.5/0.75 | better, all three: it fixes the level crudely |
| diminishing-returns | power-law trend, optionally reverting to it, or the cycle shape rescaled to it | better (both power-law variants); cycle shape on the power law inconclusive |
| vol-window | EWMA blends, shorter or longer windows | worse: the plain 365-day window wins |
| vol-reversion | volatility reverting to a long-run level or falling trend | worse |
| cycle-vol | volatility by cycle position | inconclusive (no effect) |
| tails | Student-t, or empirical horizon-level shape | Student-t worse; empirical +5% at 1 month only |
| cycle-phase | align cycles by fraction elapsed, not days | inconclusive |
Round 2 tested the power law against two refinements (the verdict rule was
unchanged, and the holdout candidates were fixed before it ran):
| Experiment | Idea | Verdict |
|---|---|---|
| powerlaw-ou | deviations from the trend revert, so uncertainty levels off; optionally plus trend-parameter uncertainty | inconclusive: +21% to +45% at 2-4 years (intervals above zero), but -0.3% and -1% at 1 month fail the "never negative" clause. Bands too narrow without parameter uncertainty, too wide with it |
| cycle-on-powerlaw | the cycle's timing, rescaled to the power-law level | inconclusive: +18% at 2 years, slightly negative at short horizons and 4 years |
### Holdout (run once, 2026-09-24)
Scored on outcomes after 2024-11-26 for the four candidates fixed in advance
(`random_walk`, `drift_rw`, `cycle`, `powerlaw`). The holdout is ~22 months
long, so at 2-4 years it is essentially one outcome period seen from several
origins (1.4-1.9 windows), and the bootstrap intervals there mean little.
| Skill vs random walk | 1mo | 3mo | 6mo | 1y | 2y | 3y | 4y |
|---|---|---|---|---|---|---|---|
| drift_rw | -0% | +1% | +5% | +14% | +31% | +53% | -65% |
| cycle | -14% | -8% | -25% | -104% | -46% | -21% | -179% |
| powerlaw | -1% | -7% | -5% | +3% | +45% | +58% | +15% |
Consistent with development: nothing beats the random walk inside a year,
`cycle` fails badly, and `powerlaw` is the best long-horizon forecast, unbiased
at 2-3 years (mean PIT 0.53-0.56) but with intervals too wide (its 80% interval
held every 2- and 3-year outcome). The holdout is now spent for these models;
a new model needs new data to be tested honestly.
### Forward test (from 2026-09-23)
New data is the only honest test left, so forecasts are now recorded before
their outcomes exist. `just snapshot` writes every tracked model's forecast
from the latest price to `data/forecasts/<date>.csv`. It refuses if the price
data is more than two days old or the date already has a snapshot, and the
files are committed so version control dates them. `just forward` scores
every recorded forecast whose target date has passed, the same way the
backtest does.
Tracked: the four registered models plus `powerlaw_ou`, `powerlaw_ou_param`
and `cycle_on_powerlaw`, which development data couldn't settle. Snapshots
store the forecasts themselves, so if a model's definition changes it gets a
new name. The routine is `just weekly` (update, snapshot, score), then commit
the new snapshot. The first 1-month outcomes arrive 2026-10-23; 1-year results
need until late 2027, and 4-year ones until 2030.
## Usage
With Nix:
```sh
nix develop
just update # fetch new daily prices from Coinbase
just backtest # score models on development data -> output/backtest/
just forecast # forecast from the latest price -> output/forecast/
just ab # run A/B experiments -> output/ab/
just weekly # new prices, a forward-test snapshot, and its scores so far
just test
just holdout # score on held-out outcomes; sparingly
```
Without Nix, any Python 3.13 with numpy, pandas 3, scipy and matplotlib works:
`python -m btcmodel --help`.
## Layout
```
btcmodel/
data.py price loading (Investing.com archive + Coinbase), dev cutoff
halving.py halving calendar, position in cycle
forecast.py Forecast (quantiles of log price), CRPS, PIT
evaluate.py walk-forward backtest and summary
experiments.py A/B tests: hypothesis, control, variants, verdict rule
forward.py forward test: record forecasts, score them as outcomes arrive
models/ drift, volatility and shape components; register models in __init__.py
plots.py fan chart, skill and calibration charts
```
A model is any object with a `name` and
`forecast(history, horizons) -> Forecast`. `history` is a date-indexed frame
holding everything up to the forecast origin and nothing after it. Other data
sources (hash rate, on-chain metrics, macro series) can join it as extra
columns. Anything published with a lag must be shifted to the date it was
actually available, or the backtest will quietly cheat.
Price data: [Investing.com](https://www.investing.com/crypto/bitcoin/historical-data)
through 2024-11-26, then Coinbase Exchange daily closes (UTC).