Composable models and A/B tests of the 2024 ideas; add powerlaw.

Models are now a Composite of drift, volatility and (optional) shape
components, so an experiment can swap one part against a fixed control.

btcmodel/experiments.py holds seven experiments built from the ideas in the
old branches (catalogued in docs/2024-ideas.md), each with its hypothesis
and source, and a verdict rule fixed before anything ran. `just ab` runs
them on development data. Results:

- Shrinking the cycle drift, and a power-law trend (plain or reverting),
  beat their controls. The power law beats the random walk by 53-63% at
  3-4 years with unbiased outcomes, so it is promoted to MODELS.
- Every alternative volatility estimate (EWMA blends, other windows,
  reversion to a level or trend) is worse than the trailing 365-day window.
  Cycle-dependent volatility, heavy tails and stretched cycle phase show no
  reliable effect.
This commit is contained in:
sam
2026-09-24 03:01:46 -07:00
parent cfc27a38de
commit b0243adf61
18 changed files with 795 additions and 155 deletions
+7 -5
View File
@@ -55,9 +55,11 @@ def backtest(
return pd.concat(rows, ignore_index=True)
def summarize(scores: pd.DataFrame, n_boot: int = 2000, seed: int = 0) -> pd.DataFrame:
def summarize(
scores: pd.DataFrame, baseline: str = BASELINE, n_boot: int = 2000, seed: int = 0
) -> pd.DataFrame:
"""
Per model and horizon: mean CRPS, skill relative to the baseline, and coverage.
Per model and horizon: mean CRPS, skill relative to `baseline`, and coverage.
Skill is 1 - CRPS / baseline CRPS (positive = better than the baseline),
with a 90% moving-block bootstrap interval over origins. Forecasts from
@@ -68,7 +70,7 @@ def summarize(scores: pd.DataFrame, n_boot: int = 2000, seed: int = 0) -> pd.Dat
rng = np.random.default_rng(seed)
rows = []
for horizon, at_h in scores.groupby("horizon"):
base = at_h[at_h.model == BASELINE].set_index("origin")["crps"].sort_index()
base = at_h[at_h.model == baseline].set_index("origin")["crps"].sort_index()
span = (base.index[-1] - base.index[0]).days + horizon
block = max(1, min(int(np.ceil(horizon / ORIGIN_STEP_DAYS)), len(base) // 2))
boot_index = _block_bootstrap_indices(len(base), block, n_boot, rng)
@@ -99,14 +101,14 @@ def _block_bootstrap_indices(n, block, n_boot, rng) -> np.ndarray:
return ((starts[:, :, None] + np.arange(block)) % n).reshape(n_boot, -1)[:, :n]
def format_summary(summary: pd.DataFrame) -> str:
def format_summary(summary: pd.DataFrame, baseline: str = BASELINE) -> str:
table = pd.DataFrame(
{
"model": summary["model"],
"horizon": summary["horizon"].map(horizon_label),
"windows": summary["windows"].map("{:.1f}".format),
"crps": summary["crps"].map("{:.3f}".format),
"skill vs rw [90%]": [
f"skill vs {baseline} [90%]": [
f"{s:+.0%} [{lo:+.0%}, {hi:+.0%}]"
for s, lo, hi in zip(summary.skill, summary.skill_lo, summary.skill_hi, strict=True)
],