Marcin Pawlowski 4baccd8d4b
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.

Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.

The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.

Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
  taken, so shared streams must give bit-identical trajectories. All 200
  replicate pairs differ by exactly 0.0. Unpaired, the same control only
  had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
  +-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
  chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
  Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
  delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
  both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
  pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
  the same quantity (t = 2.8) had failed correction.

So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.

Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
  reached the parquet; plot_fine_delay.py falls back to the unpaired
  test when it cannot confirm pairing, so the sweep would have completed
  and quietly reported the old result. Caught before the run finished;
  the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
  gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
  itself when the streams are shared, and falls back to the t-test only
  when there is real spread.

§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.

Tests: 214 passed (was 209). ruff clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
..

Total-Stake-Inference parameter selection

Per-node network simulation of Cryptarchia Total Stake Inference (TSI). Simulator: tsi-sim-pernode. All runs at the true security parameter k = 2160 unless noted; latency is in slots and 1 slot = 1 s.

This report selects and justifies the TSI parameters for Cryptarchia from a per-node network simulation. The whole report is one document — tsi-report.md — and section numbers (§1§9, Appendices AC) are stable identifiers referenced from the simulator and from spec discussion.

Uncle references. The model analysed throughout is the countable one: counting-only references, deduplicated by slot, drawn from first-fork blocks only, within a window derived as w_u = W_abs/f. An unrestricted baseline — any orphan in the window at any fork depth — is measured alongside it for comparison. The two are indistinguishable in the design regime ρ < 1 and diverge only under overload. See the model note at the top of the report, the mechanism in §2.1, the comparison in §3.2§3.2a, and the reproduction notes in §9.

Contents

Read the report →

§ what it covers
§1 executive summary — the problem, the findings, the recommendation
§2 the model, the measurement convention, and the counting rule
§3 the findings and their evidence, including the high-precision design band (§3.2a)
§4 design equations and the parameter-selection algorithm
§5 caveats and regime of validity
§6 robustness — jitter, grinding, withholding, selfish mining, rewards, reorg depth, churn
§7 parameter reference — what each knob does
§8 the safest selection, residual risks, and the recommendation-vs-spec deltas
§9 reproducibility — how to re-run every study
A · B · C the residual f-rounding offset · the per-epoch noise floor · consensus detail

Headline recommendation

Cryptarchia baseline f = 1/30. Two design choices are foundational: count uncles per occupied slot, not per block — the density-bug fix that lands the estimate at exactly D (§2.1, §8.5) — and make genesis a single protocol constant, identical at every node, never client-configurable, since a per-node divergence is never self-corrected (§8.1 row 7). The settings: security k = 2160, uncle window W = 300 slots, uncle cap U ≥ ⌈ρ⌉ + 1 (2 at the Blend target; the protocol's MAX_UNCLES = 4 sits safely above it), learning rate β = 1, on-chain f at 10⁻⁶ precision, peering degree ≥ 6 at scale, soft uncle rewards with w_u + w_n < 1, and operate at load ρ = f·D_vis < 1. The full recommended-configuration table and rationale are in §8 →.

Figures

Figures are embedded from report-figures/ via relative links and are versioned here alongside the report. They are produced by the simulator's plotting scripts (scripts/*.py and tsi_sim.plotting) in tsi-sim-pernode; that simulation folder does not commit its own generated figures — the copies checked in here are the report's figures of record.

Reproducing the results

The simulation code, configs, and run data live in tools/simulators/tsi/tsi-sim-pernode. Every study's exact command is listed in §9 — Reproducibility. In short, from the simulator directory: make install, then make <config> to run a sweep (results land under runs/<timestamp>_<label>/), and the per-figure generators under scripts/ render the figures. Regenerated figures must be copied into report-figures/ to update this report.