Marcin Pawlowski ef82ed614d
Review pass: reproduce every number from its data of record, fix what did not
Correctness/completeness review of the report and simulator. Verified against
the committed parquets: the sec 6.6 countable-ceiling table (cap-64 MDP sweep),
sec 6.10 Result 4's depth ceilings, the sec 3.4 uncle-selection table, all
adversary-variant numbers, the rho-boundary row-4 quotes (0.976 at rho=0.91,
4-sigma shortfall at 0.96, max cell 1.0024), and the sec 8.4 capstone table.
Three defects found, all fixed:

1. The collapse event was not reproducible from the committed script. Study D
   swept only the default (random) coalition, but the one observed collapse is
   a whale cell; the "once in 144 runs" count came from an ad-hoc probe. The
   committed sweep now carries the selection axis (96 runs) and reproduces the
   event: 1/12 in the whale 50% cell at delta_max = 8, never at 4. All six
   fold-related passages now quote the committed sweep, which also retires the
   stale "the full dynamics never reach it" wording in the sec 6 arc, the
   sec 6.2 intro, row 6 and item 1 -- text that contradicted item 18 since
   yesterday's finding.

2. capstone.py's printout could not reproduce the report's sec 8.4 table. The
   report's numbers are a per-replicate-tail aggregation (each replicate burns
   in against its own early-stop length); the script cut the tail at the ARM's
   max epoch, silently dropping any replicate that stopped earlier (7 of 8 in
   the adversary arm) and landing one rounding step off on three cells. The
   script now aggregates per replicate and prints the SEM; against the existing
   parquet it reproduces the table exactly (1.001/0.998, 0.342+-0.009 /
   0.343+-0.005, p_ref 1.000/0.990, 8 reps both arms). The report table was
   right all along; sec 6.8's p_ref quote (0.989, the per-arm value) is aligned
   to 0.990.

3. Small report fixes: slow-beta deflation rounded 0.765 -> "0.77" (now 0.76);
   fig13's caption now points at the fig36 ceiling instead of implying free
   recovery; row 5 cites the measured slow-beta standing deflation; the
   canonical-data paragraph lists the new studies' artifacts; the simulator
   README's layout block lists the new tests and scripts.

Adds a unit test for reorg.countable_recovery_from_depths (the one new
function that had none). 236 tests pass; the new-study parquets are copied to
the main checkout's runs/, where every other study's data of record lives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 12:29:13 +02:00

60 lines
2.7 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

#!/usr/bin/env python
"""Capstone: the recommended configuration end-to-end at true k=2160 (report §8).
One config — f=1/30, W=300, U=2, β=1, degree 6, Blend 3 hops × 8 s, Pareto stake — run honest
and under a 30 % uncle-suppression adversary, confirming accuracy, consensus, fork rate, reorg
depth, and the emergent reference rate p_ref ALL hold together. Writes runs/capstone.parquet.
"""
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src"))
import pandas as pd
from joblib import Parallel, delayed
from tsi_sim.config import SimConfig
from tsi_sim.engine import run_trajectory
REC = dict(n_nodes=1000, stake_dist="pareto", topology="blend", degree=6,
link_latency_mean=0.5, link_latency_dist="geo", blend_hops=3, blend_delay_max=8.0,
max_uncles=2, uncle_window=300, uncle_strategy="oldest", k=2160, epochs=40,
genesis_d_factor=0.5, early_stop=True)
def _one(adv: float, rep: int) -> list[dict]:
cfg = SimConfig(**REC, adversary_frac=adv, adversary_strategy="suppress", replicate=rep)
rows = run_trajectory(cfg)
for r in rows:
r["adv"] = adv
return rows
def main() -> None:
out = Path(__file__).resolve().parents[1] / "runs"
jobs = [(a, r) for a in (0.0, 0.3) for r in range(8)]
res = Parallel(n_jobs=4, backend="loky", inner_max_num_threads=1)(
delayed(_one)(a, r) for a, r in jobs)
df = pd.DataFrame([row for traj in res for row in traj])
df.to_parquet(out / "capstone.parquet", index=False)
print("=== Capstone: recommended config, all metrics together (equilibrium tail) ===")
for adv, g in df.groupby("adv"):
# Per-REPLICATE tail: early_stop ends replicates at different epochs, so a per-arm cut
# (epoch >= arm_max//2) would silently drop any replicate that stopped before the cut
# and skew the tail toward the slow-converging ones. The report's §8.4 numbers are the
# per-replicate aggregation; keep this printout matching them.
t = pd.concat([r[r.epoch >= r.epoch.max() // 2] for _, r in g.groupby("replicate")])
per_rep = t.groupby("replicate").fork_rate.mean()
sem = per_rep.std(ddof=1) / (len(per_rep) ** 0.5)
print(f"adversary {adv:.0%}: D̂/D {t.mean_ratio.mean():.4f} "
f"range_ratio {t.range_ratio.max():.4f} agreement {t.agreement_window.min():.4f} "
f"fork_rate {per_rep.mean():.3f}+-{sem:.3f}(SEM over {len(per_rep)} reps) "
f"max_reorg_depth {t.max_reorg_depth.max()} p_ref {t.p_ref.mean():.3f}")
print(f"wrote {out/'capstone.parquet'} ({len(df)} rows)")
if __name__ == "__main__":
main()