Two findings from re-reviewing the fine-delay section. 1. The rho values I put in s3.2a were wrong. The report derives rho = f*D_vis with D_vis = hops*delta_max/2 + (hops+1)*ell_mean from a MEASURED ell_mean (1.211 slots at N=1000/degree=6), not from the link_latency_mean parameter (0.5). Hand-substituting a guessed 1.5 inflated every value by ~0.04: the band is rho 0.21-0.41, not 0.25-0.45. To stop that recurring, graph_ell_mean moves out of rho_boundary_analysis.py into figures_pernode.py, joined by a new rho_for() that both scripts and any future quotation go through; plot_fine_delay.py now prints the derived rho per delay. This also exposed an inconsistency in the existing s3.2 table, which rounded delta_max=4 to "rho ~ 0.4" while s3.2a called the same cell 0.36 and prose elsewhere already used 0.56 for delta_max=8. The s3.2 column now carries the derived values (0.36/0.56/0.96/1.76). 2. Testing each cell against the exact target 1.0 -- the same question the gap test asks, without reference to the other model -- corroborates the first-fork onset independently. Unrestricted: 1/15 cells below 1 (t=-2.09, chance). Countable: 4/15, and not scattered -- delta_max=4 at U=1, and ALL THREE caps at delta_max=5 (-0.0012 to -0.0019, t=-2.5..-3.7). A shortfall appearing at every cap at once, only at the top of the band, only under the restricted model, is the first-fork cost seen absolutely. That makes "one uncle slot is sufficient -- not approximately, exactly" too strong as I had written it. s3.2a now states the residual (0.1-0.2% at the top of the band, zero below delta_max=3), reconciles it with the s1 headline, and notes that since all three caps show the same shortfall the residual is not a capacity limit. The bound quoted in s1 moves from "below 0.15%" to "<= 0.2%". Also adds the new run directories to s9's canonical list, which covered every other study but not these. Tests: 209 passed. ruff clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tsi-sim-pernode — Cryptarchia TSI per-node network simulator (Phase 2)
The reduced-model simulators (
../tsi-sim/,../tsi-sim-mc/) collapse the network to one global canonical chain and one scalarD_estper epoch. This package removes that collapse: every one of theNnodes runs TSI individually with its ownD_est, from its own partial view of the block tree under explicit message propagation over a peering graph. Its job is to test the reduced model's assumption that all honest nodes agree.
What it models
- Per-node lottery: node
iwins a slot withφ_f(w_i / D_est_i)—D_estis a length-Nvector, each node self-updating from its own view (the reduced model's key reuse: the sparse sampler already takes a per-node probability vector). - Topology (
topology): three propagation models over the network.full_mesh: every node one hop away with uniform latencyL— reproduces the reduced model exactly (validation baseline).regular: a random d-regular peering graph (configurabledegree) with per-link latency (link_latency_dist ∈ {fixed, uniform, exp, geo}, all with meanlink_latency_mean). A block reaches a node after the shortest weighted path from its producer (gossip flooding). Models direct block gossip.blend: the same d-regular graph, but a block is first routed through the Blend mixnet before it is public — the producer picksblend_hopsdistinct relay nodes uniformly at random, the block hopsproducer → r₁ → … → r_hopsover the graph, each relay waiting aUniform(0, blend_delay_max)mixing delay before forwarding, and the last relay's forward is the final network-wide gossip that makes the block visible. Relays are blind forwarders (they learn the block only from that final gossip). The dominant latency is the per-hop mixing, not the graph transport — this is the multi-slot regime where forks and the stake underestimate appear and uncle references matter. Because the mixing delays areUniform-bounded, the windowed fork choice stays exact (horizon(blend_hops+1)·max_path_latency + blend_hops·blend_delay_max).
- Real-world latency (units). Latency is in slots and a slot is 1 s, so measured
internet latencies (tens–hundreds of ms) are fractions of a slot; arrivals are therefore
kept sub-slot (float), not rounded to whole slots.
link_latency_dist=geodraws each link from a geographic band mixture (~15 msmetro →~200 msantipodal, EU↔EU ≪ EU↔AU), rescaled solink_latency_meanstays the mean-latency knob. Soregularruns the realistic sub-slot direct-gossip regime (~0.05–0.2slot), where forks are rare, andblendruns the multi-slot Blend-mixnet regime, where per-hop mixing delays dominate. - Per-node views: one global block tree plus an
(N × n_blocks)arrival matrixA; each node builds on / measures density over the blocks that have arrived at it. Uncle refs are baked at production from the producer's view (faithful — immutable once adopted). - Uncle model (
uncle_model, CLI--old): the default countable model implements the spec's counting-only rules (cryptarchia-v1-protocol.md): only the first block of a fork (parent on the producer's chain) is referenceable/countable, the window is derived asw_u = window_absorption / fslots (Wexpected block-intervals, defaultW = 10→ 300 slots, boundedW ≤ 0.6·k), selection skips slots already occupied on the producer's chain and picks one uncle per slot, and the measurement pass re-checks every rule per reference (rejections tallied asdeep_ref_share). Passing--oldtotsi-sweep/tsi-verifyruns the pre-redesign model unchanged — window =uncle_windowslots, any-depth orphans referenceable, every baked reference counted — and bit-reproduces historical runs (the old model's RNG key is byte-identical to the pre-uncle_modelkey). - Metrics: per-node
D_estspread (range,IQR), canonical-chain agreement (window prefix vs current tip), mean accuracy, and — withinit_dest=heterogeneous— transient re-convergence.
Headline result
Per-node D_est disagreement collapses to zero. Because TSI reads density from a window
buried far past k-finality, and all nodes seed the recursion from a common hardcoded
genesis D, every node computes the same measured density m → identical D_est
(range ≈ 0, agreement_window = 1) — even under a sparse graph with high latency and heavy
tip-level forking (agreement_tip can drop well below 1). This validates the reduced
model. Topology/latency instead shift the shared mean accuracy (via fork rate → q),
which uncle references recover just as in the reduced model. (Injected heterogeneous
disagreement, which the real protocol never creates, is preserved by the common
multiplicative update — a cautionary note, not protocol behaviour.)
Quick start
make install # venv + editable install
make test # unit + fast per-node checks
make verify # per-node validation (parity, spread→0, agreement, topology effect)
# Run any configs/<name>.yaml by its stem (auto-discovered); each writes a dated runs/ folder:
make smoke # tiny end-to-end grid + figures (configs/smoke.yaml)
make default # scaled-k divergence/topology sweep + figures (configs/default.yaml)
make fullscale # full-scale (true k) confirmation (configs/fullscale.yaml)
make figures RESULTS=runs/<dir>/results.parquet # re-render figures from a run
# Extra sweep flags: make fullscale SWEEP_ARGS="--batch-size 1 --mem-frac 0.6"
Scale & performance
- Representation: one global block tree +
(N × n_blocks)float64arrival matrixA(sub-slot arrivals); topologypath_latency[N,N](per-node Dijkstra, once per trajectory). n_blocksis NOT~10·kin general — it tracks block production.n_blocksis the number of lottery wins in an epoch,≈ E·Σᵢφ(wᵢ/D_est). At equilibrium that is~10·k(≈ 22k at k=2160), but whenD_estis far below the true stake — the collapsed-estimate regime, e.g. a smallgenesis_d_factor—Σ(stake)/D_est = 1/genesis_d_factoris large and block production explodes proportionally. Atgenesis_d_factor=0.01, genesis epoch-0 produces ~2.0M blocks (100× equilibrium) →A ≈ 15 GBfor a single worker;D_estself-corrects to equilibrium within ~2 epochs, so only the earliest epoch(s) are heavy. Raisinggenesis_d_factortoward 0.1–0.5 collapses this cost (0.1 → ~0.22M blocks → ~1.6 GB; 0.5 → ~44k → ~0.4 GB) and does not change the equilibrium result, which is measured after burn-in.- Sliding-window prune (
prune_arrival, default on): the arrival matrix never needs per-node columns for blocks past the horizon — under deterministic latency a block withslot ≤ t − Hhas reached every node, so its column is finalized and dropped. We keep columns only for blocks insidemax(horizon, uncle_window)slots in a base-offset buffer, turning theO(N·n_blocks)matrix intoO(N · keep-span-blocks). This is what makes the collapsed regime affordable: at N=1000/k=2160/gdf=0.01the buffer is ~tens of MB instead of the ~15 GB full matrix (fork choice, the parent clamp, uncle selection, and per-node tips all reconstruct exactly from it). It is bit-identical to the full matrix atjitter_mean == 0(proven bytest_prune_matches_full_matrixacross topologies/uncles/gdf); with jitter it falls back to the full matrix (whose safety clamp keeps the tree valid). Setprune_arrival: falseto force the full matrix (the parity oracle). The measurement pass also argmaxes in node-row bands so it adds only a small temporary. Divergence sweeps run at scaled k=256 (configs/default.yaml); full-scale k=2160 is validated to N ≤ 2000. - Worker sizing (auto, RAM-safe): both
A(~N·n_blocks, incl. the block explosion above) andpath_latency(~N²) grow, so the sweep runner sizes the loky pool to fit a RAM budget (--mem-frac, default 0.7 of physical RAM) instead of blindly using every core. The per-worker estimate realises the seeded stake to compute the genesis-epoch block count (expected_peak_blocks), so it reflects a low-genesis_d_factorexplosion rather than assuming~10·k. A calibration probe measures a real worker's peak RSS (one genesis epoch of the heaviest config in a spawned process) whenever the estimate is heavy orN > 2000(--calibrate {auto,always,never}, defaultauto; the probe bounds itself to physical RAM so it fails loud rather than freezing). - Fail-loud memory guard (
memguard.py): every worker checks size before allocating both big arrays — the(N × n_blocks)A(inbuild_tree_pernode) and the(N × N)path_latency(inbuild_path_latency, built first) — and raisesArrivalMatrixTooLargeif it would exceed the budgetTSI_ARRIVAL_BYTES_BUDGET. The sweep sets that to each worker's RAM share; unset or0is not "unlimited" — it resolves toDEFAULT_BUDGET_FRAC(0.9) of physical RAM, so a barerun_trajectory,tsi-verify, the probe, or a--mem-frac 0run all keep an absolute per-process ceiling. So a mis-estimated block explosion (or a hugeN) fails with a clear message instead of freezing the machine. - Cost: dominated by the per-node fork choice (batched per slot) and the arrival-matrix fill; the sparse lottery is negligible. Across-config joblib loky parallelism reused.
- Measurement optimisation (
measure.py): the per-node canonical/density/agreement pass was ~95% of an epoch. It is now deduped by tip (nodes sharing a tip share every derived quantity — high agreement collapsesNto a handful of computations) and the per-tip chain walk runs as a cached numba kernel (pure-Python fallback if numba is absent). Exact — bit-identical to the naive loop (test_measure). Measured ~9× end-to-end (heavy config 11.3 s → 1.2 s) and ~14× on measurement-bound configs. numba comes via theaccelextra (pip install -e ".[dev,accel]", done bymake install). - Windowed fork choice (
windowed_fork_choice, default on): bounds the block-tree build's per-slot candidate scan to a horizon of the max path latency plus the fully-propagated best tip, turningO(n_blocks^2)fork choice intoO(n_blocks*H). Exact when link latency is deterministic (jitter_mean == 0) — bit-identical to a full scan (parity test). Withjitter_mean > 0it becomes a tiny approximation and warns; a safety clamp still keeps the tree valid, andwindowed_fork_choice=Falseforces a guaranteed-exact full scan. - Reproducibility: every draw spawns off
SeedSequence(hash(config))— child 0 stake, 1 graph, 2 init, 3+e epoche.graph_seed/degree/link_latency_*are part of the config identity.- numpy-version caveat: the
accelextra (numba) requiresnumpy<2.5, so installing it pins numpy to 2.4.x. numpy'sGenerator.choice(replace=False)is not stream-stable across the 2.4↔2.5 boundary, and the sparse lottery uses it heavily only in the degenerate collapsed-estimate regime (D_est → 0⇒ win-prob → 1 ⇒count ≈ n_slots). So a run on numpy 2.5 and a run on numpy 2.4 give identical results for all normal configs but can diverge chaotically in that one extreme regime (e.g.degree=4, link_latency=8, where the estimate has already collapsed to ~0.13 — off the safe chart). The differences are tiny (max |Δ mean_ratio| ≈ 3e-3) and change no conclusion; pin numpy if bit-reproducibility across environments is required.
- numpy-version caveat: the
Layout
src/tsi_sim/ constants config rng stake lottery topology blocktree(+build_tree_pernode)
uncles(+select_uncles_at_production) tsi(+update_D_vec) epoch engine metrics
theory verify plotting/{style, figures_pernode, make_figures}
configs/ smoke.yaml default.yaml fullscale.yaml
countable-vs-old.yaml absorption-window.yaml (countable-model studies)
fine-delay.yaml (delay 1-5 at 40 replicates: the design band, high precision)
tests/ test_{pernode,config,rng,lottery,blocktree,uncles,tsi_counting,stake,
theory,latency,theory_convergence,countable_counting,...}.py
scripts/ plot_countable_vs_old.py (countable-vs-unrestricted comparison figures)
plot_fine_delay.py (design-band accuracy + model gap with 95% CIs)