Rewrite as a probabilistic model with walk-forward evaluation.

Replace the 2024 model (model.py, ~2000 lines) with the btcmodel package, the
baseline for future work:

- Forecasts are quantiles of log price at each horizon, scored with CRPS in a
  walk-forward backtest (origins every 30 days from 2014, horizons 1 month to
  4 years). Skill is relative to a zero-drift random walk, with circular
  block-bootstrap intervals and a count of independent windows.
- Development data stops at 2024-11-26, the last day the 2024 model saw.
  Later outcomes are a holdout, scored only by `backtest --holdout`.
- Models: random_walk, drift_rw, and cycle (the 2024 model's cycle-position
  drift, now kernel-smoothed and recency-weighted). On development data
  nothing beats the random walk with confidence; cycle loses at every horizon.
- Prices: the Investing.com archive moves to data/ (cut at 2024-11-26; its
  last row was intraday) and is extended with Coinbase daily closes by
  `update`.

Also: Nix flake dev shell (Python 3.13, pandas 3), ruff in place of black,
pytest suite, and a rewritten README. NOTES.md is removed as inaccurate, and
poetry is dropped.
This commit is contained in:
sam
2026-09-24 02:19:02 -07:00
parent eefff47070
commit cfc27a38de
24 changed files with 1767 additions and 3124 deletions
+34
View File
@@ -0,0 +1,34 @@
"""Halving calendar and position within the halving cycle."""
import numpy as np
import pandas as pd
GENESIS = pd.Timestamp("2009-01-03")
# Block heights 210k, 420k, 630k, 840k (UTC dates).
HALVINGS = pd.DatetimeIndex(["2012-11-28", "2016-07-09", "2020-05-11", "2024-04-20"])
# Later halvings are projected at the length of the last cycle. Block times drift
# by weeks per cycle, which is noise at the resolution this is used.
_LAST_CYCLE = HALVINGS[-1] - HALVINGS[-2]
_PROJECTED = pd.DatetimeIndex([HALVINGS[-1] + k * _LAST_CYCLE for k in range(1, 6)])
# Genesis starts cycle 0.
CYCLE_STARTS = pd.DatetimeIndex([GENESIS]).append(HALVINGS).append(_PROJECTED)
def cycle_position(dates) -> tuple[np.ndarray, np.ndarray]:
"""
For each date, return (cycle index, days since that cycle began).
Cycle 0 runs from genesis to the first halving; a halving day is day 0 of
the cycle it starts.
"""
dates = pd.DatetimeIndex(dates)
if (dates < GENESIS).any():
raise ValueError("date before genesis")
if (dates >= CYCLE_STARTS[-1]).any():
raise ValueError("date beyond projected halvings")
index = CYCLE_STARTS.searchsorted(dates, side="right") - 1
days = (dates - CYCLE_STARTS[index]).days
return np.asarray(index), np.asarray(days)