Files
bitcoin-model/docs/2024-ideas.md
T
sam b0243adf61 Composable models and A/B tests of the 2024 ideas; add powerlaw.
Models are now a Composite of drift, volatility and (optional) shape
components, so an experiment can swap one part against a fixed control.

btcmodel/experiments.py holds seven experiments built from the ideas in the
old branches (catalogued in docs/2024-ideas.md), each with its hypothesis
and source, and a verdict rule fixed before anything ran. `just ab` runs
them on development data. Results:

- Shrinking the cycle drift, and a power-law trend (plain or reverting),
  beat their controls. The power law beats the random walk by 53-63% at
  3-4 years with unbiased outcomes, so it is promoted to MODELS.
- Every alternative volatility estimate (EWMA blends, other windows,
  reversion to a level or trend) is worse than the trailing 365-day window.
  Cycle-dependent volatility, heavy tails and stretched cycle phase show no
  reliable effect.
2026-09-24 03:01:46 -07:00

3.2 KiB
Raw Blame History

Ideas in the 2024 model's history

A catalogue of the modeling ideas in the old model.py and its branches (September 2026). Labels are referenced from btcmodel/experiments.py. "Fudge" means the constant or rule existed to hit backtest coverage/MAPE targets rather than to express a hypothesis.

Drift

  • D1. Mean return by halving-cycle position. The core idea throughout (initial commit, Switch to simpler log-based projection, Implement trend smoothing, tuning-b). Smoothing went Savitzky-Golay → 60-day centred mean → Gaussian kernel with recency weights. The old code stretched every cycle to 1460 days (phase = fraction of the halving interval). Bugs: simple returns compounded as log returns (+~25%/yr bias) in the initial commit; genesis date off by a year; tuning-3 regressed log price on row index across cycles. The "blend ratio of 0 is best?" commit was blending the simple- and log-return versions of the same quantity.
  • D2. Drift damping. Constants everywhere (×0.6–0.9, asymmetric, era- and cycle-position-keyed, a 3%/day cap, the fundamentals ×0.65–0.75). Fudge as implemented; the hypothesis underneath (the cycle drift is overfit) is real.
  • D3. Diminishing returns. Initial commit decayed drift by 0.9^(t/365); old/backtests-trend-enhancement-1 regressed per-cycle returns on cycle number, but added that on top of the cycle drift (double counting) and clipped paths.
  • D4. "Skew". loc += sign(μ)·0.087σ, tuned so 68% coverage hit 68.1%. A fudge; its honest cousin is momentum.
  • D5. Stock-to-flow. A 30% blend of the daily % change in S2F into the drift. Broken: mismatched units, and S2F "halvings" on the wrong dates created a large fake post-halving drift.

Volatility

  • V1. EWMA blend. 30/90/180-day spans, weights .5/.3/.2 then .2/.5/.3, times a 1.2 fudge.
  • V2. Adaptive windows / regime weights. Short/long vol ratio rescales window lengths and blend weights; over-parameterised, and after the fundamentals rewrite the adaptive windows were computed but unused.
  • V3. Cycle-position volatility. old/backtest-vol-cycle.
  • V4. Market maturity. Volume growth, inverse vol, autocorrelation and a post-futures dummy, min-max normalised in-sample. Removed as "complexity without clear benefit".
  • V5. "Fundamentals". Supply growth, volume/supply and "depth" combined and clipped to [0.65, 0.75]: effectively a constant, grid-searched to match the era constants it replaced.
  • V6. Market conditions. Vol ratio, MA50−MA200 trend strength (≈30× too large from a units bug) and drawdown widening uncertainty.

Calibration (all fudges)

  • C1. Era scaling keyed to the backtest's training start date, with stacked factors (net ≈0.45× vol) and hand-picked era boundaries.
  • C2. Horizon uncertainty multipliers, caps and floors on daily vol.
  • C3. Changing the quantile levels: a "95%" band read off the ~81% quantiles past a year.

Shape

  • S1. Skewed innovations, with skew set per era (tuning-1).

Also

Plumbing (backtest framework, multiprocessing, output handling, cycle and CDPR plots). The old NOTES.md described "machine learning" and "macro indicator" experiments; no such code ever existed.