Files
bitcoin-model/README.md
T
sam 67d016fe25 Forward test: record forecasts before their outcomes exist.
`snapshot` writes each tracked model's forecast quantiles at the seven
backtest horizons from the latest price to data/forecasts/<origin>.csv. It
refuses stale data (older than two days) and duplicate dates, so snapshots
can't be reconstructed after the fact; committing them dates them.
`forward` scores every recorded forecast whose target date has passed,
reusing the backtest's scoring (now factored out as evaluate.score).

Tracked: random_walk, drift_rw, cycle, powerlaw, plus powerlaw_ou,
powerlaw_ou_param and cycle_on_powerlaw, which development data couldn't
settle. `just weekly` runs update, snapshot and forward.

First snapshot: 2026-09-23 (BTC $84.4K). The first outcomes are due
2026-10-23.
2026-09-24 03:08:30 -07:00

8.9 KiB
Raw Blame History

Bitcoin Price Model

I have no idea what I'm doing

Don't take this seriously. It's all in good fun.

The 2024 edition

In November 2024 I asked Claude (3.6 Sonnet, via copy and paste in Claude Web) to help build a Bitcoin price model. After a lot of branching it produced ~2000 lines of cycle analysis, "market fundamentals", era adjustments and Monte Carlo, and forecasts like this one (as of 2024-11-14):

moooooon

In September 2026 a newer Claude scored the forecasts it made on 2024-11-27 against what actually happened:

2024 forecast (median) Actual
Cycle top 2025-10-21, $199K 2025-10-06, $124.7K
2026-06-30 $133K (95% floor $65K) $58.5K
2026-09-23 $124K $84.4K

It called the timing of the top within 15 days, 19 months ahead, but put the level ~60% too high. "The price stays at $92.7K" was a better forecast (18% mean error against 68%). Its "95%" intervals were really ~81% intervals past a year out, thanks to a fudge factor that narrowed them at long horizons.

That code lives in the jj/git history. This is the rewrite.

How it works now

A forecast is a probability distribution of the price at each horizon, not a line. Every model produces quantiles of log price, and every model is scored the same way:

  • Walk-forward. Every 30 days from 2014 on, each model sees only the data up to that day and forecasts 1 month to 4 years ahead.
  • CRPS. Each forecast is scored against what happened with the continuous ranked probability score, in log-price units (0.1 ≈ "typically 10% off"). It rewards being sharp and being calibrated at once, and can't be gamed by narrowing or widening intervals.
  • Baselines. Skill is reported relative to a zero-drift random walk, with a 90% block-bootstrap interval. Nearby forecasts overlap heavily, so the report also shows windows: the number of genuinely independent outcomes.
  • Holdout. Development only sees data up to 2024-11-26, the last day the 2024 model saw. Everything after it is held out, and is scored only by just holdout, once per round of model changes. (Caveat: we already know roughly what happened in 2025-26, so it isn't perfectly blind.)

Models

  • random_walk: zero drift; "it stays about here, give or take".
  • drift_rw: drift equal to the last four years' average; "it keeps doing what it did last cycle".
  • cycle: the 2024 model's one real idea. Expected return depends on the day of the halving cycle, estimated from past cycles, with recent cycles weighted more.
  • powerlaw: log price grows linearly in log time since genesis, so growth keeps slowing. Fitted walk-forward; the exponent has stayed between 5.4 and 6.0 in every fit since 2014.

Findings so far (development data)

  • cycle loses to the random walk at every horizon, and so does every setting tried (recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days). The level is the problem: each cycle has grown less than the last (log return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any average of past cycles overshoots.
  • powerlaw models exactly that, and it is the first model to beat the random walk with some confidence: +53% and +63% skill at 3 and 4 years, with unbiased outcomes (mean PIT 0.51). Only 3-4 independent windows back that up, the functional form is famous because it fits Bitcoin's history, and the holdout hasn't been run yet.
  • Its intervals are too wide at long horizons (the 80% interval held every 3-year outcome), because it treats deviations from the trend as permanent.

A/B tests of the 2024 ideas

just ab runs the experiments in btcmodel/experiments.py: ideas salvaged from the old branches (catalogued in docs/2024-ideas.md), each a control plus variants that change one component. Hypotheses and the verdict rule were written down before anything ran. With ~100 comparisons, expect a few flukes either way.

Experiment Idea Verdict
shrink-cycle scale the cycle drift by 0.25/0.5/0.75 better, all three: it fixes the level crudely
diminishing-returns power-law trend, optionally reverting to it, or the cycle shape rescaled to it better (both power-law variants); cycle shape on the power law inconclusive
vol-window EWMA blends, shorter or longer windows worse: the plain 365-day window wins
vol-reversion volatility reverting to a long-run level or falling trend worse
cycle-vol volatility by cycle position inconclusive (no effect)
tails Student-t, or empirical horizon-level shape Student-t worse; empirical +5% at 1 month only
cycle-phase align cycles by fraction elapsed, not days inconclusive

Round 2 tested the power law against two refinements (the verdict rule was unchanged, and the holdout candidates were fixed before it ran):

Experiment Idea Verdict
powerlaw-ou deviations from the trend revert, so uncertainty levels off; optionally plus trend-parameter uncertainty inconclusive: +21% to +45% at 2-4 years (intervals above zero), but -0.3% and -1% at 1 month fail the "never negative" clause. Bands too narrow without parameter uncertainty, too wide with it
cycle-on-powerlaw the cycle's timing, rescaled to the power-law level inconclusive: +18% at 2 years, slightly negative at short horizons and 4 years

Holdout (run once, 2026-09-24)

Scored on outcomes after 2024-11-26 for the four candidates fixed in advance (random_walk, drift_rw, cycle, powerlaw). The holdout is ~22 months long, so at 2-4 years it is essentially one outcome period seen from several origins (1.4-1.9 windows), and the bootstrap intervals there mean little.

Skill vs random walk 1mo 3mo 6mo 1y 2y 3y 4y
drift_rw -0% +1% +5% +14% +31% +53% -65%
cycle -14% -8% -25% -104% -46% -21% -179%
powerlaw -1% -7% -5% +3% +45% +58% +15%

Consistent with development: nothing beats the random walk inside a year, cycle fails badly, and powerlaw is the best long-horizon forecast, unbiased at 2-3 years (mean PIT 0.53-0.56) but with intervals too wide (its 80% interval held every 2- and 3-year outcome). The holdout is now spent for these models; a new model needs new data to be tested honestly.

Forward test (from 2026-09-23)

New data is the only honest test left, so forecasts are now recorded before their outcomes exist. just snapshot writes every tracked model's forecast from the latest price to data/forecasts/<date>.csv. It refuses if the price data is more than two days old or the date already has a snapshot, and the files are committed so version control dates them. just forward scores every recorded forecast whose target date has passed, the same way the backtest does.

Tracked: the four registered models plus powerlaw_ou, powerlaw_ou_param and cycle_on_powerlaw, which development data couldn't settle. Snapshots store the forecasts themselves, so if a model's definition changes it gets a new name. The routine is just weekly (update, snapshot, score), then commit the new snapshot. The first 1-month outcomes arrive 2026-10-23; 1-year results need until late 2027, and 4-year ones until 2030.

Usage

With Nix:

nix develop
just update     # fetch new daily prices from Coinbase
just backtest   # score models on development data -> output/backtest/
just forecast   # forecast from the latest price -> output/forecast/
just ab         # run A/B experiments -> output/ab/
just weekly     # new prices, a forward-test snapshot, and its scores so far
just test
just holdout    # score on held-out outcomes; sparingly

Without Nix, any Python 3.13 with numpy, pandas 3, scipy and matplotlib works: python -m btcmodel --help.

Layout

btcmodel/
  data.py       price loading (Investing.com archive + Coinbase), dev cutoff
  halving.py    halving calendar, position in cycle
  forecast.py   Forecast (quantiles of log price), CRPS, PIT
  evaluate.py   walk-forward backtest and summary
  experiments.py  A/B tests: hypothesis, control, variants, verdict rule
  forward.py    forward test: record forecasts, score them as outcomes arrive
  models/       drift, volatility and shape components; register models in __init__.py
  plots.py      fan chart, skill and calibration charts

A model is any object with a name and forecast(history, horizons) -> Forecast. history is a date-indexed frame holding everything up to the forecast origin and nothing after it. Other data sources (hash rate, on-chain metrics, macro series) can join it as extra columns. Anything published with a lag must be shifted to the date it was actually available, or the backtest will quietly cheat.

Price data: Investing.com through 2024-11-26, then Coinbase Exchange daily closes (UTC).