diff --git a/reports/tsi/report-figures/fig34_fine_delay_accuracy.png b/reports/tsi/report-figures/fig34_fine_delay_accuracy.png
index f2dffbc..db577ff 100644
Binary files a/reports/tsi/report-figures/fig34_fine_delay_accuracy.png and b/reports/tsi/report-figures/fig34_fine_delay_accuracy.png differ
diff --git a/reports/tsi/report-figures/fig35_fine_delay_gap.png b/reports/tsi/report-figures/fig35_fine_delay_gap.png
index cc879f5..1b4401d 100644
Binary files a/reports/tsi/report-figures/fig35_fine_delay_gap.png and b/reports/tsi/report-figures/fig35_fine_delay_gap.png differ
diff --git a/reports/tsi/tsi-report.md b/reports/tsi/tsi-report.md
index 38963b1..f6bbf5f 100644
--- a/reports/tsi/tsi-report.md
+++ b/reports/tsi/tsi-report.md
@@ -21,7 +21,7 @@
---
-> **The uncle-reference model analysed here.** Uncle references are *counting-only* — they never affect a block's validity ([§6.7](#s6-7)–[§6.8](#s6-8)) — and are deduplicated by **slot**, not by block ([§8.5](#s8-5)). The reference window is derived from a **window absorption parameter**, `w_u = W_abs/f` with `W_abs = 10` expected block-intervals and bound `W_abs ≤ 0.6·k`, which is the `W = 300 = 10/f` sizing rule of [§3.4](#s3-4). One structural constraint shapes the results: only the **first block of a fork** — one whose parent lies on the referencing chain — is countable, which is the price of verifying every reference from chain data alone. This report calls that the **countable** model, and measures it against an **unrestricted** baseline in which any orphan in the window is referenceable at any fork depth ([§2.1](#s2-1), [§3.2](#s3-2), [§9](#s9)). The result: **in the design regime `ρ < 1` the two are practically identical.** A dedicated 40-replicate sweep of the operating band (`δ_max` 1–5, `ρ ≈ 0.21–0.41`) puts every `U ≥ 1` cell of both models in 0.998–1.001, resolves no difference in any individual cell (widest 95 % CI ±0.15 pp), and bounds the first-fork cost at **≤ 0.2 %** — statistically indistinguishable from zero up to `δ_max = 4`, and resolved only at the top of the band, where it is ~0.14 % ([§3.2a](#s3-2a)). Differences appear only at `ρ ≥ 1`: the restriction trims ~1.4 % at `ρ ≈ 1`, and under overload (`ρ ≈ 1.8`) recovery saturates near 0.95 where the unrestricted model reaches ~1.00. Deep forks are unrecoverable by construction, so past `ρ ≈ 1` the ceiling is set by the fork *depth* profile and **no uncle cap buys it back** — a second, independent reason for the [§8](#s8) "keep `ρ < 1`" rule. The closed-form accuracy with the measured effective utilization, `log(1−f)/log(1−f/q_u)`, predicts every countable cell to within 0.2 %.
+> **The uncle-reference model analysed here.** Uncle references are *counting-only* — they never affect a block's validity ([§6.7](#s6-7)–[§6.8](#s6-8)) — and are deduplicated by **slot**, not by block ([§8.5](#s8-5)). The reference window is derived from a **window absorption parameter**, `w_u = W_abs/f` with `W_abs = 10` expected block-intervals and bound `W_abs ≤ 0.6·k`, which is the `W = 300 = 10/f` sizing rule of [§3.4](#s3-4). One structural constraint shapes the results: only the **first block of a fork** — one whose parent lies on the referencing chain — is countable, which is the price of verifying every reference from chain data alone. This report calls that the **countable** model, and measures it against an **unrestricted** baseline in which any orphan in the window is referenceable at any fork depth ([§2.1](#s2-1), [§3.2](#s3-2), [§9](#s9)). The result: **in the design regime `ρ < 1` the two are practically identical.** A dedicated 40-replicate **paired** sweep of the operating band (`δ_max` 1–5, `ρ ≈ 0.21–0.41`), in which both arms share the stake draw, the peering graph and every lottery outcome so the uncle rule is the only difference, puts every `U ≥ 1` cell of both models in 0.998–1.001 and pins the first-fork cost at **nothing measurable up to `δ_max` = 3 and ~0.1 % at `δ_max` = 4–5** ([§3.2a](#s3-2a)). Differences appear only at `ρ ≥ 1`: the restriction trims ~1.4 % at `ρ ≈ 1`, and under overload (`ρ ≈ 1.8`) recovery saturates near 0.95 where the unrestricted model reaches ~1.00. Deep forks are unrecoverable by construction, so past `ρ ≈ 1` the ceiling is set by the fork *depth* profile and **no uncle cap buys it back** — a second, independent reason for the [§8](#s8) "keep `ρ < 1`" rule. The closed-form accuracy with the measured effective utilization, `log(1−f)/log(1−f/q_u)`, predicts every countable cell to within 0.2 %.
## 1. Executive summary
@@ -193,35 +193,33 @@ So: **at the operating loads (`ρ < 1`) no difference between the models is dete
### 3.2a The design band at high precision — `δ_max` 1–5
-**The question.** [§3.2](#s3-2) samples the mixing delay at 4/8/16/32. That resolves the overload regime, but it leaves the band the parameters are actually chosen in — the low-delay end, where every operating point sits — measured at four-fold spacing and five replicates, which is enough to say "no difference detected" and nothing more. How accurate is TSI across `δ_max` = 1–5, and how large a first-fork cost can be *excluded* there?
+**The question.** [§3.2](#s3-2) samples the mixing delay at 4/8/16/32 with five replicates. That resolves the overload regime, but it leaves the band the parameters are actually chosen in measured at four-fold spacing and low power — enough to say "no difference detected" and nothing more. How accurate is TSI across `δ_max` = 1–5, and how large a first-fork cost can be *excluded* there?
-**The finding.** One uncle slot holds the estimate at the true stake across the whole band: every `U ≥ 1` cell under both referencing models lands in **0.998–1.001**, flat in delay, while the uncle-free baseline decays 0.810 → 0.640. The two referencing models are **indistinguishable through `δ_max` ≤ 3**; at the top of the band the first-fork restriction becomes measurable and costs **0.1–0.2 %**. Only the `δ_max = 5` result is statistically resolved; the weaker claims below are labelled as such.
+**The finding.** One uncle slot holds the estimate at the true stake across the whole band: every `U ≥ 1` cell under both referencing models lands in **0.998–1.001**, flat in delay, while the uncle-free baseline decays 0.806 → 0.623. The first-fork restriction costs **nothing measurable up to `δ_max` = 3** and about **0.1 %** at `δ_max` = 4–5.
-`configs/fine-delay.yaml` spends replicates instead of range — `δ_max` ∈ {1,2,3,4,5}, `U` ∈ {0,1,2,4}, **40 replicates** per cell (8× [§3.2](#s3-2)), N = 1 000, blend, 3 hops — run under both models ([§9](#s9)). The band spans `ρ ≈ 0.21` to `0.41`, entirely inside the design regime.
+**The design is paired.** The two referencing models normally draw independent RNG streams — `uncle_model` is part of the config key, which is what lets `--old` bit-reproduce earlier runs ([§9](#s9)) — so a comparison pays the between-run variance twice and the arms differ in stake draw, peering graph and every lottery outcome. `configs/fine-delay-paired.yaml` sets `paired_streams`, which derives the RNG root from the model-independent part of the key: both arms then get the **same stake, the same graph and the same lottery draws**, and the uncle rule is the only difference. Each replicate becomes a matched pair, the shared variance cancels, and the test is a one-sample t on the 40 per-replicate differences. Trajectories still diverge after epoch 0 through the genuine feedback — a different counted density changes the next epoch's difficulty — which is the effect being measured, not noise. Grid: `δ_max` ∈ {1,2,3,4,5}, `U` ∈ {0,1,2,4}, 40 replicates, N = 1 000, blend, 3 hops, spanning `ρ ≈ 0.21` to `0.41`.
-| `δ_max` (s) | ρ | `U=0` countable / unrestricted | `U≥1` countable | `U≥1` unrestricted | pooled gap (t) |
+| `δ_max` (s) | ρ | `U=0` (both models) | `U≥1` countable | `U≥1` unrestricted | pooled gap (t) |
|---|---|---|---|---|---|
-| 1 | ≈ 0.21 | 0.810 / 0.806 | 0.9994 – 0.9999 | 0.9990 – 1.0010 | −0.0004 ± 0.0007 (1.1) |
-| 2 | ≈ 0.26 | 0.762 / 0.763 | 0.9996 – 1.0002 | 0.9996 – 1.0003 | +0.0000 ± 0.0008 (0.0) |
-| 3 | ≈ 0.31 | 0.716 / 0.723 | 0.9995 – 1.0002 | 0.9992 – 1.0004 | −0.0001 ± 0.0008 (0.2) |
-| 4 | ≈ 0.36 | 0.664 / 0.668 | 0.9988 – 0.9996 | 0.9995 – 0.9999 | −0.0005 ± 0.0008 (1.3) |
-| 5 | ≈ 0.41 | 0.640 / 0.623 | 0.9982 – 0.9988 | 0.9995 – 1.0003 | **−0.0014 ± 0.0007 (3.7)** |
+| 1 | ≈ 0.21 | 0.806 | 0.9995 – 1.0004 | 0.9990 – 1.0010 | −0.0003 ± 0.0005 (1.1) |
+| 2 | ≈ 0.26 | 0.763 | 0.9993 – 0.9996 | 0.9996 – 1.0003 | −0.0004 ± 0.0005 (1.8) |
+| 3 | ≈ 0.31 | 0.723 | 0.9993 – 0.9997 | 0.9992 – 1.0004 | −0.0003 ± 0.0004 (1.4) |
+| 4 | ≈ 0.36 | 0.668 | 0.9982 – 0.9997 | 0.9995 – 0.9999 | **−0.0011 ± 0.0004 (4.7)** |
+| 5 | ≈ 0.41 | 0.623 | 0.9986 – 0.9992 | 0.9995 – 1.0003 | **−0.0009 ± 0.0005 (3.7)** |
-**No individual cell resolves a model difference.** Across the 15 `U ≥ 1` cells the widest 95 % CI half-width is ±0.0015, one cell clears `t = 2` (0.75 are expected to by chance), and its `t = 2.59` does not survive the Bonferroni threshold of 2.94 for 15 tests. Read cell by cell, the honest statement is that any difference is **smaller than ±0.15 pp**.
+**The negative control is now an identity check.** With `U = 0` no reference is ever taken, so under shared streams the two arms must produce *bit-identical* trajectories — not merely agree within noise. They do: **all 200 replicate pairs differ by exactly 0.0**, at every delay. That is a far stronger validation than the unpaired design could offer, where the same control only had to agree within a ±0.025 band and in fact drifted by 0.016.
-**Pooled, a delay-dependent cost appears at the top of the band.** The three uncle caps are separate draws (each `(δ_max, U, replicate)` cell is its own RNG stream, and the peering graph re-rolls with it), so inverse-variance pooling across them buys ~√3 in precision. Pooled per delay, `δ_max` 1–4 stay unresolved (|t| ≤ 1.3) while **`δ_max = 5` resolves at −0.0014 ± 0.0007 (t = 3.7)**. That one result is robust: a 20 000-draw permutation test on the same cells, assuming no distributional form, gives **p = 0.0003** (all other delays: p ≥ 0.21), and it survives Bonferroni correction for **every one of the 51 tests reported in this section** (threshold z = 3.30). Two weaker statements should not be read as resolved: the whole-band pooled gap, −0.00048 ± 0.00033 (t = 2.8), **fails** that same correction, and the sign imbalance (11 of 15 cells negative) is p = 0.12 on its own. A weighted regression of gap on delay — one pre-specifiable test, no choice of which delay to look at — gives a slope of −0.00024 ± 0.00012 per slot (t = −2.05): consistent with a monotone cost, but on its own marginal. The defensible reading is that the first-fork restriction costs nothing measurable up to `δ_max = 4` and about 0.14 % at `δ_max = 5`, which is negligible against the ±0.9 % per-epoch sampling noise of [Appendix B](#sB). (Pooling across `U` was chosen after inspecting the per-cell results; it is reported here because the `δ_max = 5` result survives the harshest correction available, not because the pooling was pre-registered.)
+**What pairing bought.** Per-cell standard errors shrink by a median factor of 1.6 (range 1.2–2.1), taking the widest 95 % CI from ±0.0015 to **±0.0010**. Five of the fifteen `U ≥ 1` cells now resolve at |t| ≥ 2 against 0.75 expected by chance, and the largest (`U = 2`, `δ_max = 4`, t = 4.32) clears the Bonferroni threshold of 2.94 for 15 tests. The unpaired sweep resolved only `δ_max = 5`, and only after pooling.
-**The same data, measured against 1.0 rather than against the other arm.** Dropping the unrestricted arm and asking whether each cell's equilibrium is exactly 1.0 removes one noisy measurement, so it is ~1.4× more sensitive. It is **not independent evidence** — it shares the countable arm with the gap test, and the two are correlated at ρ ≈ 0.73 — but the *contrast between the arms* is informative. Under the **unrestricted** model 1 of 15 cells sits below 1 (`t` = −2.09), consistent with chance. Under the **countable** model 4 of 15 do, and they are not scattered: `δ_max = 4` at `U = 1` (−0.0012, `t` = −2.6) and **all three uncle caps at `δ_max = 5`** (−0.0012 to −0.0019, `t` = −2.5 to −3.7; the `U = 2` cell survives the 51-test correction). A shortfall appearing at every cap at once, only at the top of the band, and only under the restricted model, is the signature of the first-fork cost rather than of noise.
+**The cost is a step, not a ramp.** Pooling across the three uncle caps — separate draws of the same underlying difference — `δ_max` 1–3 stay unresolved (t = 1.1, 1.8, 1.4) and then **`δ_max` = 4 and 5 both resolve**, at −0.0011 (t = 4.7) and −0.0009 (t = 3.7). Over the whole band the pooled gap is −0.00060 ± 0.00021, **t = 5.7** — comfortably resolved, where the unpaired estimate of the same quantity (t = 2.8) did not survive correction. So the restriction switches on somewhere around `δ_max` ≈ 4 rather than growing smoothly with delay, and above that it sits near 0.1 % without deepening across the band tested. Measured against the exact target 1.0 instead of against the other arm — a more sensitive view of the same countable data, not independent evidence — the same picture appears: the unrestricted model has 1 of 15 cells below 1 (chance), the countable model has 5, all at `δ_max` 4–5.
-**What limits the resolution: the comparison is unpaired.** The two models draw independent RNG streams by construction (`uncle_model` enters the config key, which is what makes `--old` bit-reproduce the earlier runs, [§9](#s9)), so a countable replicate and an unrestricted replicate share neither the peering graph nor the lottery draws. Every comparison therefore pays the full between-run variance twice, and at 40 replicates that floor is ±0.0015 per cell — an order of magnitude above the effect being measured. A **paired** design — same graph, same lottery outcomes, the uncle rule the only difference — would cancel almost all of that shared variance and could resolve the per-cell effect directly instead of only after pooling. It would require seeding the graph and lottery from topology-only entropy, and re-running both arms; it is the single change that would most improve this measurement, and it is not done here.
+**How big is 0.1 %.** An order of magnitude below the ±0.9 % per-epoch sampling noise of [Appendix B](#sB), and below the ~1 % fixed-point rounding bias of [Appendix A](#sA). The [§1](#s1) statement that one uncle restores the estimate to the true stake holds at the precision that matters; this section puts the residual at ~0.1 % at the top of the band rather than at zero.
-**The negative control passes.** The `U = 0` arms — identical models by construction — show |gap| ≤ 0.016 with max `t` = 1.26, i.e. no spurious signal, but a 95 % CI of ±0.025: **17× wider than the entire `U ≥ 1` range.** The unrecovered regime is intrinsically noisy, which is precisely why the model comparison has to be made where uncles are active.
+
-
+
-
-
-**What this settles.** One uncle slot is sufficient everywhere in the operating band under either referencing model — the [§3.3](#s3-3) `U ≥ ⌈ρ⌉` rule has margin to spare at `ρ ≤ 0.41`, and raising `U` past 1 buys nothing here (at `δ_max = 5` all three caps sit at the same 0.1–0.2 % shortfall, so the residual is *not* a capacity limit). And the first-fork restriction, which [§3.2](#s3-2) shows costing 1.4 % at `ρ ≈ 1` and 5 % under overload, costs **at most 0.2 %** anywhere a deployment should be operating — an order of magnitude below the ±0.9 % per-epoch sampling noise of [Appendix B](#sB), and below the ~1 % fixed-point rounding bias of [Appendix A](#sA). The [§1](#s1) statement that one uncle restores the estimate to the true stake holds at the precision that matters; this section puts the residual at 0.1–0.2 % at the top of the band rather than zero.
+**What this settles.** One uncle slot is sufficient everywhere in the operating band under either referencing model — the [§3.3](#s3-3) `U ≥ ⌈ρ⌉` rule has margin to spare at `ρ ≤ 0.41`, and raising `U` past 1 does not remove the residual (all three caps show the same ~0.1 % shortfall at `δ_max` = 4–5, so it is not a capacity limit). And the first-fork restriction, which [§3.2](#s3-2) shows costing 1.4 % at `ρ ≈ 1` and 5 % under overload, costs **about 0.1 %** at the top of the design band and nothing measurable below it.
### 3.3 One uncle is not always enough — the load `ρ`
@@ -845,14 +843,15 @@ Sweep studies are committed configs, run with `make ` (writes a dated `run
| U = 0 fluctuation series (zero delay, k ∈ {256, 1024, 2160}) | `scripts/appendix_fluct.py --run` (`runs/fluctuation_u0.parquet`) | [Appendix B](#sB) |
| CI smoke grid + analytic sanity checks | `configs/smoke.yaml`; `scripts/verify.py` (`make verify`) — validation only, no figures | — |
| **countable vs unrestricted referencing** (accuracy over delay × U; measured `q_u`/recovery `r`) | `configs/countable-vs-old.yaml` run twice — default and with `--old`; figures + significance table by `scripts/plot_countable_vs_old.py` | [§2.1](#s2-1), [§3.2](#s3-2) |
-| **fine delay band** (δ_max 1–5 at 40 replicates; tight CI on the model gap in the design regime) | `configs/fine-delay.yaml` run twice — default and with `--old`; figures + significance tables (model gap, and each model vs the exact 1.0 target) by `scripts/plot_fine_delay.py` | [§3.2a](#s3-2a) |
+| **fine delay band, PAIRED** (δ_max 1–5, 40 replicates, common random numbers — the primary design-regime measurement) | `configs/fine-delay-paired.yaml` run twice — default and with `--old`; paired-difference tests + figures by `scripts/plot_fine_delay.py` | [§3.2a](#s3-2a) |
+| fine delay band, unpaired (same grid without `paired_streams`; superseded, retained as the power comparison) | `configs/fine-delay.yaml` run twice | [§3.2a](#s3-2a) |
| **window absorption sweep** (`W` in expected block-intervals, `w_u = W/f` derived) | `configs/absorption-window.yaml`; figure by `scripts/plot_countable_vs_old.py` | [§3.4](#s3-4) |
**Uncle-model convention.** The simulator's default is the **countable** model — first-fork candidates only, derived window `w_u = W/f`, occupied-slot exclusion, per-reference counting rules ([§2.1](#s2-1)). The **unrestricted** baseline is preserved in the code and selected with `--old` on `tsi-sweep`/`tsi-verify`. Its RNG key is byte-identical to the pre-restriction key, so `--old` **bit-reproduces the earlier runs**: a `rho-boundary` cell (δ_max = 8, U = 2, k = 256, N = 1 000) re-run under `--old` matches the committed `2026-07-27_195627_rho-boundary` parquet with `max |Δ| = 0` on every epoch and every metric. Studies in the table above that predate the countable default were produced under the unrestricted model and reproduce exactly under `--old`; the comparison rows quantify where the two models differ, and in the design regime (`ρ < 1`) no difference is resolvable, so those findings carry over unchanged.
Because the two models draw independent RNG streams, every countable-vs-unrestricted comparison is **unpaired**, and its resolution is set by the replicate spread rather than by the effect size. Each comparison sweep therefore includes a `U = 0` arm as a **negative control**: with no uncles the models are identical by construction, so the measured `U = 0` gap is a direct reading of the noise floor at that delay and replicate count. At `δ_max = 32` with 5 replicates that floor is ≈ 0.23 in `D̂/D` — larger than several real effects elsewhere in the grid — which is why [§3.2](#s3-2) reports a `t` statistic per cell and why the design regime is measured separately at 40 replicates ([§3.2a](#s3-2a)).
-All studies were **re-run on 2026-07-23/24 with the corrected slot-counting mechanism** ([§2.1](#s2-1)) and the early-stop optimisation; the resilient batch is `scripts/run_all_reruns.sh` (per-step log in `runs/rerun_status.log`). Canonical run directories (latest): fullscale N=5000/10000 = `2026-07-24_094519_fullscale`; fullscale N=1000/2000 = `2026-07-23_171803_fullscale-small`; uncle-window = `2026-07-24_001456`; window-uncles = `2026-07-24_014240`; block-rate = `2026-07-24_043943`; blend-hops-delay = `2026-07-24_064052`; window-scale = `2026-07-24_085234`; latency-shape = `2026-07-24_090014_expdist`; stake-tail = `2026-07-24_090044_pareto133`; heterogeneous-start = `2026-07-24_090114_default`; N-scaling = the `nscaling-{a,b}` + `nscaling32-{a,b}` runs; adversary grids = `runs/adversary_grid/`; jitter = `runs/jitter_grid/`; bootstrap = `runs/bootstrap_fullscale/`; fluctuation = `runs/fluctuation_u0.parquet`; fork-rate = `runs/fork_rate_vs_delay.parquet`; ρ-boundary = `2026-07-27_195627_rho-boundary`. Referencing-model studies: countable-vs-unrestricted = `2026-08-04_103536_cvo-countable` / `2026-08-04_104010_cvo-old`; window absorption = `2026-08-04_104633_absorption-window`; fine delay band = `2026-08-04_191441_fine-countable` / `2026-08-04_195353_fine-old`.
+All studies were **re-run on 2026-07-23/24 with the corrected slot-counting mechanism** ([§2.1](#s2-1)) and the early-stop optimisation; the resilient batch is `scripts/run_all_reruns.sh` (per-step log in `runs/rerun_status.log`). Canonical run directories (latest): fullscale N=5000/10000 = `2026-07-24_094519_fullscale`; fullscale N=1000/2000 = `2026-07-23_171803_fullscale-small`; uncle-window = `2026-07-24_001456`; window-uncles = `2026-07-24_014240`; block-rate = `2026-07-24_043943`; blend-hops-delay = `2026-07-24_064052`; window-scale = `2026-07-24_085234`; latency-shape = `2026-07-24_090014_expdist`; stake-tail = `2026-07-24_090044_pareto133`; heterogeneous-start = `2026-07-24_090114_default`; N-scaling = the `nscaling-{a,b}` + `nscaling32-{a,b}` runs; adversary grids = `runs/adversary_grid/`; jitter = `runs/jitter_grid/`; bootstrap = `runs/bootstrap_fullscale/`; fluctuation = `runs/fluctuation_u0.parquet`; fork-rate = `runs/fork_rate_vs_delay.parquet`; ρ-boundary = `2026-07-27_195627_rho-boundary`. Referencing-model studies: countable-vs-unrestricted = `2026-08-04_103536_cvo-countable` / `2026-08-04_104010_cvo-old`; window absorption = `2026-08-04_104633_absorption-window`; fine delay band (unpaired) = `2026-08-04_191441_fine-countable` / `2026-08-04_195353_fine-old`; fine delay band (paired, primary) = `2026-08-05_111109_fine-paired-countable` / `2026-08-05_113218_fine-paired-old`.
Figures are in `report-figures/` (`fig1`–`fig29`, plus [Appendix B](#sB)'s `figB1`–`figB2`; numbering is generation order, not order of appearance). Committed generators: `fig1` (bootstrap, k=2160) by `scripts/bootstrap_dynamics.py`; `fig2`,`fig4`,`fig5`,`fig17`–`fig22` by `scripts/regenerate_extra_figs.py` from the latest sweeps (`fig3` hops×delay×U grid by `scripts/hops_delay_grid.py`, `fig6` (block-rate `U_min` grid + ρ-collapse) rendered ad hoc from `runs/2026-07-24_043943_block-rate` with no committed generator; `fig26` deficit-vs-ρ by `scripts/rho_boundary_analysis.py`) (fullscale-derived `fig17`–`fig20` pool both sizes in that run, N = 5 000 and N = 10 000 — the generators filter on stake_dist/topology/degree/init_dest only, never on `n_nodes`); `fig8`,`fig9` by `scripts/adversary_figs.py` from `runs/adversary_grid/`; `fig10`–`fig12` by `scripts/dynamic_withhold.py`; `fig13`–`fig15` by `scripts/selfish_mining.py`/`selfish_rewards.py`/`reward_mandate.py`; `fig16` by `scripts/stake_vs_delay.py`; `fig23`–`fig24` by `scripts/nscaling_analysis.py`; `fig25` by `scripts/window_scale_analysis.py`; `fig27`–`fig28` by `scripts/reorg_depth.py` (fork rates via `--measure`; private-chain model `src/tsi_sim/reorg.py`); `fig29` by `scripts/churn.py`; `figB1`–`figB2` by `scripts/appendix_fluct.py`. `fig7` (feedback fixed-point) is an analytic overlay; `fig30`–`fig33` (countable-vs-unrestricted accuracy, `q_u`-prediction check, recovery rate, absorption-window sweep) by `scripts/plot_countable_vs_old.py` from the `cvo-countable`/`cvo-old`/`absorption-window` runs; `fig34`–`fig35` (design-regime accuracy and the model gap with 95 % CIs) by `scripts/plot_fine_delay.py` from the `fine-countable`/`fine-old` runs. Every figure type the per-node simulator generates appears in this report, and the fork-rate/reorg-depth study closes the previous reproducibility gap for the adversarial figures (`fig8`,`fig9` now have committed generators from `runs/adversary_grid/`).
diff --git a/tools/simulators/tsi/tsi-sim-pernode/configs/fine-delay-paired.yaml b/tools/simulators/tsi/tsi-sim-pernode/configs/fine-delay-paired.yaml
new file mode 100644
index 0000000..255fc79
--- /dev/null
+++ b/tools/simulators/tsi/tsi-sim-pernode/configs/fine-delay-paired.yaml
@@ -0,0 +1,38 @@
+# PAIRED high-precision sweep of the low mixing-delay band (the design regime).
+#
+# Identical grid to fine-delay.yaml, but with paired_streams: true. The unpaired sweep pays the
+# between-run variance twice — the two uncle models draw independent stake, graph and lottery
+# streams — which puts a ~+-0.0015 per-cell floor on the countable-vs-unrestricted comparison,
+# an order of magnitude above the effect. Only delta_max = 5 resolved there, and only after
+# pooling across uncle caps.
+#
+# paired_streams drops the uncle-model marker from the RNG root, so a countable cell and its
+# --old twin get the SAME stake draw, the SAME peering graph and the SAME lottery outcomes;
+# the uncle rule is then the only difference and the per-replicate difference is a PAIRED
+# observation. The shared variance cancels, so the paired test should resolve the per-cell
+# effect directly. Trajectories still diverge after epoch 0 through the genuine feedback (a
+# different counted density changes the next epoch's difficulty) — that is the signal.
+#
+# Run TWICE — default (countable) and with --old — same grid. U=0 remains the negative control:
+# with no uncles AND a shared stream the two arms must now agree EXACTLY, which is a far
+# stronger check than the unpaired version (there it only had to agree within noise).
+# Latency is in SLOTS (1 slot = 1 s).
+n_nodes: [1000] # network size
+stake_dist: [pareto] # heavy-tailed (realistic) stake distribution
+topology: [blend] # Blend mixnet (the multi-slot fork regime)
+degree: [6] # peering degree of the d-regular graph
+link_latency_mean: [0.5] # natural geographic transport (sub-slot)
+link_latency_dist: [geo] # real-world geographic band mixture
+blend_hops: [3] # fixed hop count; delay is the swept knob
+blend_delay_max: [1.0, 2.0, 3.0, 4.0, 5.0] # THE fine region (integer slots)
+max_uncles: [0, 1, 2, 4] # U: 0 = negative control, then the levers
+uncle_strategy: [oldest] # spec selection: oldest-first fill
+init_dest: [common] # per-node initial D_est from agreement
+replicates: 40 # matches fine-delay.yaml for a like-for-like read
+base: # per-run settings shared by every cell
+ k: 2160 # true security parameter
+ epochs: 20 # equilibrium within ~2 epochs; burn 50%
+ f: 0.03333333333333333 # slot activation coefficient (1/30)
+ genesis_d_factor: 0.5 # start near true stake (cheap epoch 0)
+ early_stop: true
+ paired_streams: true # common random numbers across the two arms
diff --git a/tools/simulators/tsi/tsi-sim-pernode/scripts/plot_fine_delay.py b/tools/simulators/tsi/tsi-sim-pernode/scripts/plot_fine_delay.py
index a2a10e9..3750fb5 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/scripts/plot_fine_delay.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/scripts/plot_fine_delay.py
@@ -46,6 +46,34 @@ def _load(run_dir: str | Path) -> pd.DataFrame:
return pd.read_parquet(Path(run_dir) / "results.parquet")
+def paired_gaps(cnt_raw: pd.DataFrame, old_raw: pd.DataFrame) -> pd.DataFrame | None:
+ """Per-replicate differences, when both arms were run with ``paired_streams``.
+
+ Under common random numbers replicate *i* of each arm shares the stake draw, the peering
+ graph and the lottery outcomes, so ``d_i = countable_i - unrestricted_i`` is a PAIRED
+ observation and the shared variance cancels. The test is then a one-sample t on the d_i,
+ which is what makes a sub-0.1 % effect reachable per cell instead of only after pooling.
+
+ Returns None when the runs are not paired, so the caller falls back to the unpaired test.
+ """
+ if not (cnt_raw.get("paired_streams", pd.Series([False])).all()
+ and old_raw.get("paired_streams", pd.Series([False])).all()):
+ return None
+ c, o = equilibrium(cnt_raw), equilibrium(old_raw)
+ keys = [DELAY, "max_uncles", "replicate"]
+ m = c[[*keys, "mean_ratio"]].merge(o[[*keys, "mean_ratio"]], on=keys,
+ suffixes=("_c", "_o"))
+ m["d"] = m.mean_ratio_c - m.mean_ratio_o
+ rows = []
+ for (dl, u), s in m.groupby([DELAY, "max_uncles"]):
+ d = s.d.to_numpy()
+ se = sem(d)
+ rows.append({DELAY: dl, "max_uncles": u, "gap": float(d.mean()), "se": se,
+ "ci95": Z95 * se, "t": abs(d.mean()) / se if se > 0 else np.inf,
+ "n_pair": len(d), "n_zero": int((d == 0.0).sum())})
+ return pd.DataFrame(rows).sort_values(["max_uncles", DELAY])
+
+
def _cells(df: pd.DataFrame) -> pd.DataFrame:
"""Per (delay, U): replicate mean, SEM and count of the equilibrium accuracy."""
return equilibrium(df).groupby([DELAY, "max_uncles"], as_index=False).agg(
@@ -146,10 +174,14 @@ def fig_gap(g: pd.DataFrame) -> plt.Figure:
span = float(np.abs(np.r_[g[g.max_uncles > 0].gap + g[g.max_uncles > 0].ci95,
g[g.max_uncles > 0].gap - g[g.max_uncles > 0].ci95]).max())
ax.set_ylim(-1.35 * span, 1.35 * span)
- ax.text(0.015, 0.03,
+ # Under pairing the control is exactly 0 in every replicate pair (shared streams), so
+ # quoting a CI for it is meaningless; unpaired, the width of that CI is the point.
+ note = ("U=0 negative control: exactly 0.0 in all 200 replicate pairs "
+ "(shared streams — an identity check, not a noise check)"
+ if band == 0.0 else
f"U=0 negative control (true gap = 0): 95% CI ±{band:.4f}, "
- f"{band / span:.0f}× outside this range",
- transform=ax.transAxes, fontsize=6.5, alpha=0.75)
+ f"{band / span:.0f}× outside this range")
+ ax.text(0.015, 0.03, note, transform=ax.transAxes, fontsize=6.5, alpha=0.75)
ax.legend(fontsize="x-small", ncol=2)
return fig
@@ -164,10 +196,17 @@ def main() -> None:
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)
- cnt_raw = _load(args.countable)
- cnt, old = _cells(cnt_raw), _cells(_load(args.old))
- g = gaps(cnt, old)
- prov = "tsi-sim-pernode fine-delay.yaml (+--old)"
+ cnt_raw, old_raw = _load(args.countable), _load(args.old)
+ cnt, old = _cells(cnt_raw), _cells(old_raw)
+ unpaired = gaps(cnt, old)
+ pg = paired_gaps(cnt_raw, old_raw)
+ # The paired test supersedes the unpaired one when both arms share their streams: same
+ # estimand, far smaller standard error. Keep the unpaired numbers for the variance-
+ # reduction report below.
+ g = unpaired if pg is None else unpaired.drop(columns=["gap", "se", "ci95", "t"]).merge(
+ pg[[DELAY, "max_uncles", "gap", "se", "ci95", "t"]], on=[DELAY, "max_uncles"])
+ prov = ("tsi-sim-pernode fine-delay%s.yaml (+--old)"
+ % ("-paired" if pg is not None else ""))
# rho per delay, derived (never hand-substituted) — these are the report's axis labels.
delays = sorted(cnt[DELAY].unique())
@@ -210,11 +249,35 @@ def main() -> None:
alle = float(np.sqrt(1.0 / w.sum()))
print(f" whole band: {allp:+.5f} +-{Z95 * alle:.5f} t={abs(allp) / alle:.2f}")
+ if pg is not None:
+ print("\n=== PAIRED design (common random numbers) ===")
+ # Variance reduction actually achieved, per cell, vs the unpaired standard error.
+ cmp = unpaired[[DELAY, "max_uncles", "se"]].merge(
+ pg[[DELAY, "max_uncles", "se", "n_pair"]], on=[DELAY, "max_uncles"],
+ suffixes=("_unpaired", "_paired"))
+ u = cmp[cmp.max_uncles > 0]
+ ratio = (u.se_unpaired / u.se_paired.replace(0, np.nan))
+ print(f" SE shrink at U>=1: median {ratio.median():.1f}x, range "
+ f"{ratio.min():.1f}-{ratio.max():.1f}x ({int(u.n_pair.min())} pairs/cell)")
+ pctl = pg[pg.max_uncles == 0]
+ exact, tot = int(pctl.n_zero.sum()), int(pctl.n_pair.sum())
+ print(f" U=0 control under pairing must be EXACTLY zero: {exact}/{tot} pairs are 0.0"
+ f" -> {'PASSES' if exact == tot else 'FAILS — streams are not shared'}")
+ res = pg[(pg.max_uncles > 0) & (pg.t >= 2)]
+ print(f" per-cell resolved at |t|>=2: {len(res)}/{len(pg[pg.max_uncles>0])}"
+ f" (unpaired: {int((unpaired[unpaired.max_uncles>0].t>=2).sum())}/15)")
+
ctl = g[g.max_uncles == 0]
if len(ctl):
- print(f"\nU=0 negative control (true gap = 0): |gap| up to {ctl.gap.abs().max():.4f}, "
- f"max t = {ctl.t.max():.2f}, 95% CI +-{ctl.ci95.max():.4f} "
- f"-> control {'PASSES' if ctl.t.max() < 2 else 'FAILS'}")
+ # Under pairing the control gap is EXACTLY 0, so its se is 0 and t is 0/0. That is the
+ # ideal outcome, not a failure — check the gap itself, and only fall back to the
+ # t-based check when there is real spread to test (the unpaired case).
+ worst = float(ctl.gap.abs().max())
+ exact = bool((ctl.se == 0).all()) if "se" in ctl else False
+ ok = (worst == 0.0) if exact else (float(ctl.t.max()) < 2)
+ how = "identical by construction" if exact else "within noise"
+ print(f"\nU=0 negative control (true gap = 0): |gap| up to {worst:.4g} ({how})"
+ f" -> control {'PASSES' if ok else 'FAILS'}")
# Absolute test: does uncle recovery actually land on 1.0? Same question as the gap
# test, asked without reference to the other model.
diff --git a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/config.py b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/config.py
index 4f4951c..3a4bd92 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/config.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/config.py
@@ -99,6 +99,19 @@ class SimConfig:
# Old model only (--old): the uncle reference window w_u in slots, set directly.
# Ignored by the countable model, which derives the window from window_absorption.
uncle_window: int = constants.W_DEFAULT
+ # COMMON RANDOM NUMBERS for countable-vs-old comparisons. Off by default, and deliberately
+ # NOT part of key() — with it off every seed is byte-identical to before, so historical runs
+ # and --old bit-reproduction are untouched.
+ #
+ # The two uncle models normally draw independent streams (uncle_model is in the key), so a
+ # comparison pays the full between-run variance TWICE and the arms differ in stake draw,
+ # peering graph and every lottery outcome. With paired_streams=True the RNG root is derived
+ # from the model-independent part of the key instead, so both arms get the SAME stake, the
+ # SAME graph and the SAME lottery draws; the only difference is the uncle rule, and the
+ # per-replicate difference becomes a paired observation with the shared variance cancelled.
+ # Trajectories still diverge legitimately after epoch 0 — a different counted density feeds
+ # back into the next epoch's difficulty — which is the effect being measured, not noise.
+ paired_streams: bool = False
max_uncles: int = 0 # U (0 = baseline, no uncles)
uncle_strategy: UncleStrategy = "oldest"
# Coin-flip inclusion prob for the "random" strategy. Only 0.5 reproduces the spec's
@@ -319,16 +332,9 @@ class SimConfig:
def period_T(self) -> int:
return constants.period_T(self.k, self.f)
- def key(self) -> tuple:
- """Hashable identity used to seed the RNG deterministically.
-
- Must include EVERY field that affects the run (guarded by test_rng), otherwise two
- distinct configs would share an RNG stream. ``uncle_model`` /
- ``window_absorption`` are appended ONLY for the countable model: an ``--old`` run's
- key is then byte-identical to the pre-redesign key, so ``--old`` bit-reproduces
- historical runs (the two models still get distinct streams from the marker).
- """
- base = (
+ def _base_key(self) -> tuple:
+ """Identity fields shared by both uncle models — see ``key`` and ``seed_key``."""
+ return (
self.n_nodes, self.stake_dist, self.pareto_shape, self.uniform_random,
self.total_stake, self.latency, self.latency_stochastic, self.uncle_window,
self.max_uncles, self.uncle_strategy, self.uncle_random_p, self.f, self.beta,
@@ -345,9 +351,29 @@ class SimConfig:
# NOTE: windowed_fork_choice and prune_arrival are deliberately excluded — they are pure
# compute/memory optimisations that consume no RNG and (at jitter_mean == 0) change no
# result, so pruned and full-matrix runs must share a seed (see test_pernode parity).
+
+ def key(self) -> tuple:
+ """Hashable identity used to seed the RNG deterministically.
+
+ Must include EVERY field that affects the run (guarded by test_rng), otherwise two
+ distinct configs would share an RNG stream. ``uncle_model`` /
+ ``window_absorption`` are appended ONLY for the countable model: an ``--old`` run's
+ key is then byte-identical to the pre-redesign key, so ``--old`` bit-reproduces
+ historical runs (the two models still get distinct streams from the marker).
+ """
if self.uncle_model == "old":
- return base # historical (pre-uncle_model) key: --old bit-compat
- return base + (self.uncle_model, self.window_absorption)
+ return self._base_key() # historical (pre-uncle_model) key: --old bit-compat
+ return self._base_key() + (self.uncle_model, self.window_absorption)
+
+ def seed_key(self) -> tuple:
+ """The identity the RNG root is actually derived from (see ``rng.seedseq_for``).
+
+ Identical to ``key`` except under ``paired_streams``, where it deliberately drops the
+ uncle-model marker so that a countable run and an ``--old`` run of the SAME cell draw
+ the SAME root seed — common random numbers, which is what makes the two arms a
+ *paired* sample (see ``paired_streams``).
+ """
+ return self._base_key() if self.paired_streams else self.key()
# Axes that can be swept; every SimConfig field is legal here.
diff --git a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/metrics.py b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/metrics.py
index 79587f2..7a977f4 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/metrics.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/metrics.py
@@ -15,6 +15,10 @@ _CONFIG_FIELDS = (
"link_latency_mean", "link_latency_dist", "blend_hops", "blend_delay_max",
"init_dest", "init_spread", "uncle_model", "window_absorption",
"uncle_window", "max_uncles", "uncle_strategy",
+ # Recorded so downstream analysis can TELL whether a countable/--old pair actually shared
+ # its RNG streams. The paired test is only valid on paired runs, and without this column
+ # the analysis silently falls back to the much weaker unpaired test.
+ "paired_streams",
"f", "beta", "k", "genesis_d_factor", "epochs", "fixed_point", "legacy_block_count",
"replicate",
"adversary_frac", "adversary_strategy", "adversary_period", "adversary_withhold_epochs",
diff --git a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/rng.py b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/rng.py
index 96316da..b586b5c 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/rng.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/src/tsi_sim/rng.py
@@ -17,7 +17,9 @@ from .config import SimConfig
def _entropy(config: SimConfig) -> int:
- payload = repr((config.root_seed, config.key())).encode()
+ # seed_key() == key() unless paired_streams is set, in which case the uncle-model marker is
+ # dropped so both arms of a comparison share one stream (common random numbers).
+ payload = repr((config.root_seed, config.seed_key())).encode()
digest = hashlib.blake2b(payload, digest_size=16).digest()
return int.from_bytes(digest, "big")
diff --git a/tools/simulators/tsi/tsi-sim-pernode/tests/test_config.py b/tools/simulators/tsi/tsi-sim-pernode/tests/test_config.py
index 26e47fa..f868711 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/tests/test_config.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/tests/test_config.py
@@ -140,8 +140,11 @@ def test_key_covers_every_field():
# uncle_window is read ONLY by the old model; under the (default) countable model it is
# an ignored field, deliberately left in the base tuple at its old position so that an
# --old run's key stays byte-identical to historical keys.
+ # paired_streams must NOT be in key(): it selects WHICH key the RNG root is derived from
+ # (see seed_key), so putting it in key() would perturb every historical seed and break
+ # --old bit-reproduction. Its own behaviour is pinned in test_rng.py.
ignored = {"root_seed", "windowed_fork_choice", "prune_arrival", "early_stop",
- "uncle_window"}
+ "uncle_window", "paired_streams"}
names = {f.name for f in dataclasses.fields(SimConfig)} - ignored
a = SimConfig()
for name in names:
@@ -152,6 +155,10 @@ def test_key_covers_every_field():
# ... and uncle_window IS distinguished under the old model, where it is live.
old = SimConfig(uncle_model="old")
assert old.key() != dataclasses.replace(old, uncle_window=old.uncle_window + 1).key()
+ # paired_streams leaves key() untouched but DOES change the seed derived from it.
+ a_paired = dataclasses.replace(a, paired_streams=True)
+ assert a.key() == a_paired.key()
+ assert a.seed_key() != a_paired.seed_key()
def test_old_model_key_is_historical():
diff --git a/tools/simulators/tsi/tsi-sim-pernode/tests/test_rng.py b/tools/simulators/tsi/tsi-sim-pernode/tests/test_rng.py
index 0c019c0..8f305fd 100644
--- a/tools/simulators/tsi/tsi-sim-pernode/tests/test_rng.py
+++ b/tools/simulators/tsi/tsi-sim-pernode/tests/test_rng.py
@@ -32,3 +32,77 @@ def test_replicates_differ():
b = run_trajectory(SimConfig(n_nodes=300, topology="regular", k=8, epochs=6,
link_latency_mean=2.0, replicate=1))
assert a[-1]["mean_ratio"] != b[-1]["mean_ratio"]
+
+
+def test_paired_streams_shares_the_root_across_uncle_models():
+ """Common random numbers: with paired_streams the two arms draw the SAME root seed."""
+ from tsi_sim.rng import seedseq_for
+
+ kw = dict(n_nodes=50, max_uncles=2, blend_delay_max=5.0, topology="blend",
+ k=32, epochs=2, replicate=3, paired_streams=True)
+ c = SimConfig(uncle_model="countable", **kw)
+ o = SimConfig(uncle_model="old", **kw)
+ assert c.seed_key() == o.seed_key() # the marker is dropped
+ assert c.key() != o.key() # ...but identity still distinguishes them
+ assert seedseq_for(c).entropy == seedseq_for(o).entropy
+
+
+def test_unpaired_is_the_default_and_separates_the_models():
+ kw = dict(n_nodes=50, max_uncles=2, blend_delay_max=5.0, topology="blend",
+ k=32, epochs=2, replicate=3)
+ c, o = SimConfig(uncle_model="countable", **kw), SimConfig(uncle_model="old", **kw)
+ assert c.paired_streams is False and o.paired_streams is False
+ assert c.seed_key() == c.key() and o.seed_key() == o.key()
+ assert seedseq_for(c).entropy != seedseq_for(o).entropy
+
+
+def test_paired_streams_does_not_perturb_unpaired_seeds():
+ """The flag must not enter key(): every historical seed stays byte-identical.
+
+ This is what protects --old bit-reproduction of the pre-redesign runs (report §9).
+ """
+ from tsi_sim.rng import seedseq_for
+
+ for model in ("countable", "old"):
+ base = SimConfig(uncle_model=model, n_nodes=50, max_uncles=2, k=32, epochs=2)
+ flagged = SimConfig(uncle_model=model, n_nodes=50, max_uncles=2, k=32, epochs=2,
+ paired_streams=False)
+ assert base.key() == flagged.key()
+ assert seedseq_for(base).entropy == seedseq_for(flagged).entropy
+ # and the old model's key is still exactly the base tuple (no marker appended)
+ o = SimConfig(uncle_model="old", n_nodes=50, k=32, epochs=2)
+ assert o.key() == o._base_key()
+
+
+def test_paired_streams_gives_both_arms_the_same_stake_and_graph():
+ """Pairing must reach the actual shared inputs, not just the root seed."""
+ import numpy as np
+
+ from tsi_sim import topology
+ from tsi_sim.rng import seedseq_for
+ from tsi_sim.stake import make_stake
+
+ kw = dict(n_nodes=60, degree=4, topology="blend", max_uncles=2, k=32, epochs=2,
+ blend_delay_max=5.0, replicate=1)
+ c = SimConfig(uncle_model="countable", paired_streams=True, **kw)
+ o = SimConfig(uncle_model="old", paired_streams=True, **kw)
+ kids = {n: seedseq_for(cfg).spawn(cfg.epochs + 3)
+ for n, cfg in (("c", c), ("o", o))}
+ s_c = make_stake(c, np.random.default_rng(kids["c"][0]))
+ s_o = make_stake(o, np.random.default_rng(kids["o"][0]))
+ np.testing.assert_array_equal(s_c, s_o) # same stake draw
+ g_c = topology.build_path_latency(c, np.random.default_rng(kids["c"][1]))
+ g_o = topology.build_path_latency(o, np.random.default_rng(kids["o"][1]))
+ np.testing.assert_array_equal(g_c, g_o) # same peering graph
+
+
+def test_paired_streams_is_recorded_in_the_output_row():
+ """A paired run must be identifiable from its parquet alone.
+
+ scripts/plot_fine_delay.py picks the paired test only when both arms report
+ paired_streams; if the flag were missing from the recorded config it would silently fall
+ back to the unpaired test and quietly discard the whole point of the paired sweep.
+ """
+ from tsi_sim.metrics import _CONFIG_FIELDS
+
+ assert "paired_streams" in _CONFIG_FIELDS