Files
sam b0243adf61 Composable models and A/B tests of the 2024 ideas; add powerlaw.
Models are now a Composite of drift, volatility and (optional) shape
components, so an experiment can swap one part against a fixed control.

btcmodel/experiments.py holds seven experiments built from the ideas in the
old branches (catalogued in docs/2024-ideas.md), each with its hypothesis
and source, and a verdict rule fixed before anything ran. `just ab` runs
them on development data. Results:

- Shrinking the cycle drift, and a power-law trend (plain or reverting),
  beat their controls. The power law beats the random walk by 53-63% at
  3-4 years with unbiased outcomes, so it is promoted to MODELS.
- Every alternative volatility estimate (EWMA blends, other windows,
  reversion to a level or trend) is worse than the trailing 365-day window.
  Cycle-dependent volatility, heavy tails and stretched cycle phase show no
  reliable effect.
2026-09-24 03:01:46 -07:00

66 lines
3.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ideas in the 2024 model's history
A catalogue of the modeling ideas in the old `model.py` and its branches
(September 2026). Labels are referenced from `btcmodel/experiments.py`.
"Fudge" means the constant or rule existed to hit backtest coverage/MAPE
targets rather than to express a hypothesis.
## Drift
- **D1. Mean return by halving-cycle position.** The core idea throughout
(initial commit, `Switch to simpler log-based projection`, `Implement trend
smoothing`, `tuning-b`). Smoothing went Savitzky-Golay → 60-day centred mean
→ Gaussian kernel with recency weights. The old code stretched every cycle
to 1460 days (phase = fraction of the halving interval). Bugs: simple
returns compounded as log returns (+~25%/yr bias) in the initial commit;
genesis date off by a year; `tuning-3` regressed log price on row index
across cycles. The "blend ratio of 0 is best?" commit was blending the
simple- and log-return versions of the same quantity.
- **D2. Drift damping.** Constants everywhere (×0.6–0.9, asymmetric, era- and
cycle-position-keyed, a 3%/day cap, the fundamentals ×0.65–0.75). Fudge as
implemented; the hypothesis underneath (the cycle drift is overfit) is real.
- **D3. Diminishing returns.** Initial commit decayed drift by 0.9^(t/365);
`old/backtests-trend-enhancement-1` regressed per-cycle returns on cycle
number, but added that on top of the cycle drift (double counting) and
clipped paths.
- **D4. "Skew".** `loc += sign(μ)·0.087σ`, tuned so 68% coverage hit 68.1%.
A fudge; its honest cousin is momentum.
- **D5. Stock-to-flow.** A 30% blend of the daily % change in S2F into the
drift. Broken: mismatched units, and S2F "halvings" on the wrong dates
created a large fake post-halving drift.
## Volatility
- **V1. EWMA blend.** 30/90/180-day spans, weights .5/.3/.2 then .2/.5/.3,
times a 1.2 fudge.
- **V2. Adaptive windows / regime weights.** Short/long vol ratio rescales
window lengths and blend weights; over-parameterised, and after the
fundamentals rewrite the adaptive windows were computed but unused.
- **V3. Cycle-position volatility.** `old/backtest-vol-cycle`.
- **V4. Market maturity.** Volume growth, inverse vol, autocorrelation and a
post-futures dummy, min-max normalised in-sample. Removed as "complexity
without clear benefit".
- **V5. "Fundamentals".** Supply growth, volume/supply and "depth" combined
and clipped to [0.65, 0.75]: effectively a constant, grid-searched to match
the era constants it replaced.
- **V6. Market conditions.** Vol ratio, MA50−MA200 trend strength (≈30× too
large from a units bug) and drawdown widening uncertainty.
## Calibration (all fudges)
- **C1. Era scaling keyed to the backtest's training start date**, with stacked
factors (net ≈0.45× vol) and hand-picked era boundaries.
- **C2. Horizon uncertainty multipliers**, caps and floors on daily vol.
- **C3. Changing the quantile levels**: a "95%" band read off the ~81%
quantiles past a year.
## Shape
- **S1. Skewed innovations**, with skew set per era (`tuning-1`).
## Also
Plumbing (backtest framework, multiprocessing, output handling, cycle and
CDPR plots). The old NOTES.md described "machine learning" and "macro
indicator" experiments; no such code ever existed.