Composable models and A/B tests of the 2024 ideas; add powerlaw.
Models are now a Composite of drift, volatility and (optional) shape components, so an experiment can swap one part against a fixed control. btcmodel/experiments.py holds seven experiments built from the ideas in the old branches (catalogued in docs/2024-ideas.md), each with its hypothesis and source, and a verdict rule fixed before anything ran. `just ab` runs them on development data. Results: - Shrinking the cycle drift, and a power-law trend (plain or reverting), beat their controls. The power law beats the random walk by 53-63% at 3-4 years with unbiased outcomes, so it is promoted to MODELS. - Every alternative volatility estimate (EWMA blends, other windows, reversion to a level or trend) is worse than the trailing 365-day window. Cycle-dependent volatility, heavy tails and stretched cycle phase show no reliable effect.
This commit is contained in:
@@ -55,9 +55,11 @@ def backtest(
|
||||
return pd.concat(rows, ignore_index=True)
|
||||
|
||||
|
||||
def summarize(scores: pd.DataFrame, n_boot: int = 2000, seed: int = 0) -> pd.DataFrame:
|
||||
def summarize(
|
||||
scores: pd.DataFrame, baseline: str = BASELINE, n_boot: int = 2000, seed: int = 0
|
||||
) -> pd.DataFrame:
|
||||
"""
|
||||
Per model and horizon: mean CRPS, skill relative to the baseline, and coverage.
|
||||
Per model and horizon: mean CRPS, skill relative to `baseline`, and coverage.
|
||||
|
||||
Skill is 1 - CRPS / baseline CRPS (positive = better than the baseline),
|
||||
with a 90% moving-block bootstrap interval over origins. Forecasts from
|
||||
@@ -68,7 +70,7 @@ def summarize(scores: pd.DataFrame, n_boot: int = 2000, seed: int = 0) -> pd.Dat
|
||||
rng = np.random.default_rng(seed)
|
||||
rows = []
|
||||
for horizon, at_h in scores.groupby("horizon"):
|
||||
base = at_h[at_h.model == BASELINE].set_index("origin")["crps"].sort_index()
|
||||
base = at_h[at_h.model == baseline].set_index("origin")["crps"].sort_index()
|
||||
span = (base.index[-1] - base.index[0]).days + horizon
|
||||
block = max(1, min(int(np.ceil(horizon / ORIGIN_STEP_DAYS)), len(base) // 2))
|
||||
boot_index = _block_bootstrap_indices(len(base), block, n_boot, rng)
|
||||
@@ -99,14 +101,14 @@ def _block_bootstrap_indices(n, block, n_boot, rng) -> np.ndarray:
|
||||
return ((starts[:, :, None] + np.arange(block)) % n).reshape(n_boot, -1)[:, :n]
|
||||
|
||||
|
||||
def format_summary(summary: pd.DataFrame) -> str:
|
||||
def format_summary(summary: pd.DataFrame, baseline: str = BASELINE) -> str:
|
||||
table = pd.DataFrame(
|
||||
{
|
||||
"model": summary["model"],
|
||||
"horizon": summary["horizon"].map(horizon_label),
|
||||
"windows": summary["windows"].map("{:.1f}".format),
|
||||
"crps": summary["crps"].map("{:.3f}".format),
|
||||
"skill vs rw [90%]": [
|
||||
f"skill vs {baseline} [90%]": [
|
||||
f"{s:+.0%} [{lo:+.0%}, {hi:+.0%}]"
|
||||
for s, lo, hi in zip(summary.skill, summary.skill_lo, summary.skill_hi, strict=True)
|
||||
],
|
||||
|
||||
Reference in New Issue
Block a user