Rewrite as a probabilistic model with walk-forward evaluation.
Replace the 2024 model (model.py, ~2000 lines) with the btcmodel package, the baseline for future work: - Forecasts are quantiles of log price at each horizon, scored with CRPS in a walk-forward backtest (origins every 30 days from 2014, horizons 1 month to 4 years). Skill is relative to a zero-drift random walk, with circular block-bootstrap intervals and a count of independent windows. - Development data stops at 2024-11-26, the last day the 2024 model saw. Later outcomes are a holdout, scored only by `backtest --holdout`. - Models: random_walk, drift_rw, and cycle (the 2024 model's cycle-position drift, now kernel-smoothed and recency-weighted). On development data nothing beats the random walk with confidence; cycle loses at every horizon. - Prices: the Investing.com archive moves to data/ (cut at 2024-11-26; its last row was intraday) and is extended with Coinbase daily closes by `update`. Also: Nix flake dev shell (Python 3.13, pandas 3), ruff in place of black, pytest suite, and a rewritten README. NOTES.md is removed as inaccurate, and poetry is dropped.
This commit is contained in:
@@ -6,18 +6,108 @@
|
||||
|
||||
**Don't take this seriously. It's all in good fun.**
|
||||
|
||||
I decided to have fun and ask Anthropic's Claude AI (3.6 Sonnet) to help build
|
||||
a Bitcoin price model, using data from [Investing.com](https://www.investing.com/crypto/bitcoin/historical-data).
|
||||
## The 2024 edition
|
||||
|
||||
The model is still a mess: lots of redundant code as we went through various
|
||||
methods for projecting future prices; the plots' colors don't render correctly.
|
||||
In November 2024 I asked Claude (3.6 Sonnet, via copy and paste in Claude Web)
|
||||
to help build a Bitcoin price model. After a lot of branching it produced ~2000
|
||||
lines of cycle analysis, "market fundamentals", era adjustments and Monte Carlo,
|
||||
and forecasts like this one (as of 2024-11-14):
|
||||
|
||||
I feel like I'm dangerous enough to know what to ask for out of a model, but
|
||||
not knowledgeable enough evaluate whether what Claude produced actually makes
|
||||
any sense. (Stats class in college was a long time ago...)
|
||||

|
||||
|
||||
## tl;dr show me the projection!
|
||||
In September 2026 a newer Claude scored the forecasts it made on 2024-11-27
|
||||
against what actually happened:
|
||||
|
||||
As of writing (2024-11-14), here is what the model generates:
|
||||
| | 2024 forecast (median) | Actual |
|
||||
|---|---|---|
|
||||
| Cycle top | 2025-10-21, $199K | 2025-10-06, $124.7K |
|
||||
| 2026-06-30 | $133K (95% floor $65K) | $58.5K |
|
||||
| 2026-09-23 | $124K | $84.4K |
|
||||
|
||||

|
||||
It called the timing of the top within 15 days, 19 months ahead, but put the
|
||||
level ~60% too high. "The price stays at $92.7K" was a better forecast (18%
|
||||
mean error against 68%). Its "95%" intervals were really ~81% intervals past a
|
||||
year out, thanks to a fudge factor that narrowed them at long horizons.
|
||||
|
||||
That code lives in the jj/git history. This is the rewrite.
|
||||
|
||||
## How it works now
|
||||
|
||||
A forecast is a probability distribution of the price at each horizon, not a
|
||||
line. Every model produces quantiles of log price, and every model is scored
|
||||
the same way:
|
||||
|
||||
- **Walk-forward.** Every 30 days from 2014 on, each model sees only the data up
|
||||
to that day and forecasts 1 month to 4 years ahead.
|
||||
- **CRPS.** Each forecast is scored against what happened with the continuous
|
||||
ranked probability score, in log-price units (0.1 ≈ "typically 10% off"). It
|
||||
rewards being sharp and being calibrated at once, and can't be gamed by
|
||||
narrowing or widening intervals.
|
||||
- **Baselines.** Skill is reported relative to a zero-drift random walk, with a
|
||||
90% block-bootstrap interval. Nearby forecasts overlap heavily, so the report
|
||||
also shows `windows`: the number of genuinely independent outcomes.
|
||||
- **Holdout.** Development only sees data up to 2024-11-26, the last day the
|
||||
2024 model saw. Everything after it is held out, and is scored only by
|
||||
`just holdout`, once per round of model changes. (Caveat: we already know
|
||||
roughly what happened in 2025-26, so it isn't perfectly blind.)
|
||||
|
||||
### Models
|
||||
|
||||
- `random_walk`: zero drift; "it stays about here, give or take".
|
||||
- `drift_rw`: drift equal to the last four years' average; "it keeps doing what
|
||||
it did last cycle".
|
||||
- `cycle`: the 2024 model's one real idea. Expected return depends on the day of
|
||||
the halving cycle, estimated from past cycles, with recent cycles weighted
|
||||
more.
|
||||
|
||||
### Findings so far (development data)
|
||||
|
||||
- Nothing beats the random walk with any confidence at any horizon. At 2-4
|
||||
years there are only 3-5 independent outcomes in the whole history.
|
||||
- `drift_rw` leads at 2-4 years (+13-18% skill, but the intervals span zero).
|
||||
- `cycle` loses to both at every horizon, and so does every setting tried
|
||||
(recency half-life 0.25-2 cycles, smoothing bandwidth 15-60 days). The cycle
|
||||
*shape* costs accuracy. The level is the problem: each cycle has grown less
|
||||
than the last (log return 4.0, 2.6, 2.0, i.e. roughly ×55, ×13, ×7), so any
|
||||
average of past cycles overshoots.
|
||||
|
||||
That points at diminishing returns as the structure worth modelling, e.g. a
|
||||
power-law trend, which is next.
|
||||
|
||||
## Usage
|
||||
|
||||
With Nix:
|
||||
|
||||
```sh
|
||||
nix develop
|
||||
just update # fetch new daily prices from Coinbase
|
||||
just backtest # score models on development data -> output/backtest/
|
||||
just forecast # forecast from the latest price -> output/forecast/
|
||||
just test
|
||||
just holdout # score on held-out outcomes; sparingly
|
||||
```
|
||||
|
||||
Without Nix, any Python 3.13 with numpy, pandas 3, scipy and matplotlib works:
|
||||
`python -m btcmodel --help`.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
btcmodel/
|
||||
data.py price loading (Investing.com archive + Coinbase), dev cutoff
|
||||
halving.py halving calendar, position in cycle
|
||||
forecast.py Forecast (quantiles of log price), CRPS, PIT
|
||||
evaluate.py walk-forward backtest and summary
|
||||
models/ one file per model family; register new ones in __init__.py
|
||||
plots.py fan chart, skill and calibration charts
|
||||
```
|
||||
|
||||
A model is any object with a `name` and
|
||||
`forecast(history, horizons) -> Forecast`. `history` is a date-indexed frame
|
||||
holding everything up to the forecast origin and nothing after it. Other data
|
||||
sources (hash rate, on-chain metrics, macro series) can join it as extra
|
||||
columns. Anything published with a lag must be shifted to the date it was
|
||||
actually available, or the backtest will quietly cheat.
|
||||
|
||||
Price data: [Investing.com](https://www.investing.com/crypto/bitcoin/historical-data)
|
||||
through 2024-11-26, then Coinbase Exchange daily closes (UTC).
|
||||
|
||||
Reference in New Issue
Block a user