The spec bounds an uncle's own slot (0 < sl_A - sl_U <= w_u) but leaves its PARENT unconstrained beyond lying on the referencing chain. So a block minted NOW, built on a chain block from arbitrarily far back, is a legal first-fork uncle: recent by its own slot, ancient by its parent's. Verifying it means deriving the epoch state and ledger root as of that ancient parent, per reference, and those are precisely the inputs the counting rules require -- so the work cannot be amortised. It costs the adversary nothing beyond lottery wins it already has; it just builds them somewhere useless. Measured with a deep_parent coalition. At the deployed operating point a 30% adversary moves the MEDIAN counted reference's reach from 54 slots back to 20,144, and the worst case to 76,778 -- the epoch boundary, ~21 hours of history, ~256x the nominal window. It is not a tail effect. The fix is a SUBSTITUTION, not an additional rule. A block strictly postdates its parent and a referenced uncle strictly precedes its referencer, so sl_A - sl_U < sl_A - sl_parent(U) <= w_u: bounding the parent bounds the uncle for free, and a both-windows variant would be identical to the parent one. Both invariants are pinned in a new test_slot_ordering.py rather than argued -- the user asked to confirm sl_A > sl_U explicitly, and it turns out to be load-bearing for the whole implication, so it is tested at three geometries plus a hand-built counting case. Under the parent anchor the same coalition reaches 292/300/300 slots at delta_max 4/8/16 -- capped by construction. Honest recovery is unaffected: 0.9993 -> 0.9999, 0.9969 -> 0.9986, 0.9791 -> 0.9858, no loss anywhere within one to two SEM, because a latency orphan's parent is recent by construction. One finding that sharpens the case: at delta_max = 16 the HONEST uncle-anchored arm already reaches 315 slots, past its own w_u = 300. Under the current rule w_u is not a bound on validation reach even with no adversary present. It only becomes a state-retention bound once anchored to the parent. Recorded as sec 6.12 with fig38, a new row in the sec 8.5 spec deltas, both new knobs in sec 7, and the study in sec 9. uncle_window_anchor and the deep_parent strategy are appended to the RNG key only when non-default, so no committed run is reseeded. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
tsi-sim-pernode — Cryptarchia TSI per-node network simulator (Phase 2)
The reduced-model simulators (
../tsi-sim/,../tsi-sim-mc/) collapse the network to one global canonical chain and one scalarD_estper epoch. This package removes that collapse: every one of theNnodes runs TSI individually with its ownD_est, from its own partial view of the block tree under explicit message propagation over a peering graph. Its job is to test the reduced model's assumption that all honest nodes agree.
What it models
- Per-node lottery: node
iwins a slot withφ_f(w_i / D_est_i)—D_estis a length-Nvector, each node self-updating from its own view (the reduced model's key reuse: the sparse sampler already takes a per-node probability vector). - Topology (
topology): three propagation models over the network.full_mesh: every node one hop away with uniform latencyL— reproduces the reduced model exactly (validation baseline).regular: a random d-regular peering graph (configurabledegree) with per-link latency (link_latency_dist ∈ {fixed, uniform, exp, geo}, all with meanlink_latency_mean). A block reaches a node after the shortest weighted path from its producer (gossip flooding). Models direct block gossip.blend: the same d-regular graph, but a block is first routed through the Blend mixnet before it is public — the producer picksblend_hopsdistinct relay nodes uniformly at random, the block hopsproducer → r₁ → … → r_hopsover the graph, each relay waiting aUniform(0, blend_delay_max)mixing delay before forwarding, and the last relay's forward is the final network-wide gossip that makes the block visible. Relays are blind forwarders (they learn the block only from that final gossip). The dominant latency is the per-hop mixing, not the graph transport — this is the multi-slot regime where forks and the stake underestimate appear and uncle references matter. Because the mixing delays areUniform-bounded, the windowed fork choice stays exact (horizon(blend_hops+1)·max_path_latency + blend_hops·blend_delay_max).
- Real-world latency (units). Latency is in slots and a slot is 1 s, so measured
internet latencies (tens–hundreds of ms) are fractions of a slot; arrivals are therefore
kept sub-slot (float), not rounded to whole slots.
link_latency_dist=geodraws each link from a geographic band mixture (~15 msmetro →~200 msantipodal, EU↔EU ≪ EU↔AU), rescaled solink_latency_meanstays the mean-latency knob. Soregularruns the realistic sub-slot direct-gossip regime (~0.05–0.2slot), where forks are rare, andblendruns the multi-slot Blend-mixnet regime, where per-hop mixing delays dominate. - Per-node views: one global block tree plus an
(N × n_blocks)arrival matrixA; each node builds on / measures density over the blocks that have arrived at it. Uncle refs are baked at production from the producer's view (faithful — immutable once adopted). - Uncle model (
uncle_model, CLI--old): the default countable model implements the spec's counting-only rules (cryptarchia-v1-protocol.md): only the first block of a fork (parent on the producer's chain) is referenceable/countable, the window is derived asw_u = window_absorption / fslots (Wexpected block-intervals, defaultW = 10→ 300 slots, boundedW ≤ 0.6·k), selection skips slots already occupied on the producer's chain and picks one uncle per slot, and the measurement pass re-checks every rule per reference (rejections tallied asdeep_ref_share). Passing--oldtotsi-sweep/tsi-verifyruns the pre-redesign model unchanged — window =uncle_windowslots, any-depth orphans referenceable, every baked reference counted — and bit-reproduces historical runs (the old model's RNG key is byte-identical to the pre-uncle_modelkey). - Adversaries (
adversary_strategy, over a coalition holdingadversary_fracof stake, selected at random oradversary_selection: whalefor the largest holders at matched stake):suppress— produces normally but references no uncles, starving the recovered density.withhold— never gossips its blocks; abstention, a dead loss to the attacker.selfish— mines a private chain and releases it under Eyal–Sirer SM1 rules, orphaning honest work. Only visibility is modelled: a coalition member's fork choice builds on the private tip whenever it leads, so the chain forms and is abandoned emergently. Forces the exact full scan, since a hidden block breaks the windowed horizon's premise.
- Metrics: per-node
D_estspread (range,IQR), canonical-chain agreement (window prefix vs current tip), mean accuracy, fork structure (fork_rate, reorg depth,p_refandp_ref_honest, anddeep_orphan_share— the share of orphans below their fork's first block, i.e. unreferenceable by construction), and — withinit_dest=heterogeneous— transient re-convergence.
Headline result
Per-node D_est disagreement collapses to zero. Because TSI reads density from a window
buried far past k-finality, and all nodes seed the recursion from a common hardcoded
genesis D, every node computes the same measured density m → identical D_est
(range ≈ 0, agreement_window = 1) — even under a sparse graph with high latency and heavy
tip-level forking (agreement_tip can drop well below 1). This validates the reduced
model. Topology/latency instead shift the shared mean accuracy (via fork rate → q),
which uncle references recover just as in the reduced model. (Injected heterogeneous
disagreement, which the real protocol never creates, is preserved by the common
multiplicative update — a cautionary note, not protocol behaviour.)
Quick start
make install # venv + editable install
make test # unit + fast per-node checks
make verify # per-node validation (parity, spread→0, agreement, topology effect)
# Run any configs/<name>.yaml by its stem (auto-discovered); each writes a dated runs/ folder:
make smoke # tiny end-to-end grid + figures (configs/smoke.yaml)
make default # scaled-k divergence/topology sweep + figures (configs/default.yaml)
make fullscale # full-scale (true k) confirmation (configs/fullscale.yaml)
make figures RESULTS=runs/<dir>/results.parquet # re-render figures from a run
# Extra sweep flags: make fullscale SWEEP_ARGS="--batch-size 1 --mem-frac 0.6"
Scale & performance
- Representation: one global block tree +
(N × n_blocks)float64arrival matrixA(sub-slot arrivals); topologypath_latency[N,N](per-node Dijkstra, once per trajectory). n_blocksis NOT~10·kin general — it tracks block production.n_blocksis the number of lottery wins in an epoch,≈ E·Σᵢφ(wᵢ/D_est). At equilibrium that is~10·k(≈ 22k at k=2160), but whenD_estis far below the true stake — the collapsed-estimate regime, e.g. a smallgenesis_d_factor—Σ(stake)/D_est = 1/genesis_d_factoris large and block production explodes proportionally. Atgenesis_d_factor=0.01, genesis epoch-0 produces ~2.0M blocks (100× equilibrium) →A ≈ 15 GBfor a single worker;D_estself-corrects to equilibrium within ~2 epochs, so only the earliest epoch(s) are heavy. Raisinggenesis_d_factortoward 0.1–0.5 collapses this cost (0.1 → ~0.22M blocks → ~1.6 GB; 0.5 → ~44k → ~0.4 GB) and does not change the equilibrium result, which is measured after burn-in.- Sliding-window prune (
prune_arrival, default on): the arrival matrix never needs per-node columns for blocks past the horizon — under deterministic latency a block withslot ≤ t − Hhas reached every node, so its column is finalized and dropped. We keep columns only for blocks insidemax(horizon, uncle_window)slots in a base-offset buffer, turning theO(N·n_blocks)matrix intoO(N · keep-span-blocks). This is what makes the collapsed regime affordable: at N=1000/k=2160/gdf=0.01the buffer is ~tens of MB instead of the ~15 GB full matrix (fork choice, the parent clamp, uncle selection, and per-node tips all reconstruct exactly from it). It is bit-identical to the full matrix atjitter_mean == 0(proven bytest_prune_matches_full_matrixacross topologies/uncles/gdf); with jitter it falls back to the full matrix (whose safety clamp keeps the tree valid). Setprune_arrival: falseto force the full matrix (the parity oracle). The measurement pass also argmaxes in node-row bands so it adds only a small temporary. Divergence sweeps run at scaled k=256 (configs/default.yaml); full-scale k=2160 is validated to N ≤ 2000. - Worker sizing (auto, RAM-safe): both
A(~N·n_blocks, incl. the block explosion above) andpath_latency(~N²) grow, so the sweep runner sizes the loky pool to fit a RAM budget (--mem-frac, default 0.7 of physical RAM) instead of blindly using every core. The per-worker estimate realises the seeded stake to compute the genesis-epoch block count (expected_peak_blocks), so it reflects a low-genesis_d_factorexplosion rather than assuming~10·k. A calibration probe measures a real worker's peak RSS (one genesis epoch of the heaviest config in a spawned process) whenever the estimate is heavy orN > 2000(--calibrate {auto,always,never}, defaultauto; the probe bounds itself to physical RAM so it fails loud rather than freezing). - Fail-loud memory guard (
memguard.py): every worker checks size before allocating both big arrays — the(N × n_blocks)A(inbuild_tree_pernode) and the(N × N)path_latency(inbuild_path_latency, built first) — and raisesArrivalMatrixTooLargeif it would exceed the budgetTSI_ARRIVAL_BYTES_BUDGET. The sweep sets that to each worker's RAM share; unset or0is not "unlimited" — it resolves toDEFAULT_BUDGET_FRAC(0.9) of physical RAM, so a barerun_trajectory,tsi-verify, the probe, or a--mem-frac 0run all keep an absolute per-process ceiling. So a mis-estimated block explosion (or a hugeN) fails with a clear message instead of freezing the machine. - Cost: dominated by the per-node fork choice (batched per slot) and the arrival-matrix fill; the sparse lottery is negligible. Across-config joblib loky parallelism reused.
- Measurement optimisation (
measure.py): the per-node canonical/density/agreement pass was ~95% of an epoch. It is now deduped by tip (nodes sharing a tip share every derived quantity — high agreement collapsesNto a handful of computations) and the per-tip chain walk runs as a cached numba kernel (pure-Python fallback if numba is absent). Exact — bit-identical to the naive loop (test_measure). Measured ~9× end-to-end (heavy config 11.3 s → 1.2 s) and ~14× on measurement-bound configs. numba comes via theaccelextra (pip install -e ".[dev,accel]", done bymake install). - Windowed fork choice (
windowed_fork_choice, default on): bounds the block-tree build's per-slot candidate scan to a horizon of the max path latency plus the fully-propagated best tip, turningO(n_blocks^2)fork choice intoO(n_blocks*H). Exact when link latency is deterministic (jitter_mean == 0) — bit-identical to a full scan (parity test). Withjitter_mean > 0it becomes a tiny approximation and warns; a safety clamp still keeps the tree valid, andwindowed_fork_choice=Falseforces a guaranteed-exact full scan. - Reproducibility: every draw spawns off
SeedSequence(hash(config))— child 0 stake, 1 graph, 2 init, 3+e epoche.graph_seed/degree/link_latency_*are part of the config identity.- numpy-version caveat: the
accelextra (numba) requiresnumpy<2.5, so installing it pins numpy to 2.4.x. numpy'sGenerator.choice(replace=False)is not stream-stable across the 2.4↔2.5 boundary, and the sparse lottery uses it heavily only in the degenerate collapsed-estimate regime (D_est → 0⇒ win-prob → 1 ⇒count ≈ n_slots). So a run on numpy 2.5 and a run on numpy 2.4 give identical results for all normal configs but can diverge chaotically in that one extreme regime (e.g.degree=4, link_latency=8, where the estimate has already collapsed to ~0.13 — off the safe chart). The differences are tiny (max |Δ mean_ratio| ≈ 3e-3) and change no conclusion; pin numpy if bit-reproducibility across environments is required.
- numpy-version caveat: the
Layout
src/tsi_sim/ constants config rng stake lottery topology blocktree(+build_tree_pernode)
uncles(+select_uncles_at_production) tsi(+update_D_vec) epoch engine metrics
theory verify plotting/{style, figures_pernode, make_figures}
configs/ smoke.yaml default.yaml fullscale.yaml
countable-vs-old.yaml absorption-window.yaml (countable-model studies)
fine-delay.yaml (delay 1-5 at 40 replicates: the design band, high precision)
uncle-selection.yaml (the spec's oldest-first rule vs a deviating proposer)
spec-point-{n5000,window,jitter}.yaml (the DEPLOYED operating point: delta_max=4
from the spec's Blend profile — size, window and per-recipient-variance arms)
tests/ test_{pernode,config,rng,lottery,blocktree,uncles,tsi_counting,stake,
theory,latency,theory_convergence,countable_counting,
countable_selfish,...}.py
scripts/ plot_countable_vs_old.py (countable-vs-unrestricted comparison figures)
plot_fine_delay.py (design-band accuracy + model gap with 95% CIs)
countable_selfish.py (first-fork ceiling under the selfish MDP; fig36)
adversary_variants.py (whale/jitter/slow-beta variants + the withhold-load sweep)
deflation_frontier.py (how far a PAID adversary can deflate D_est; fig37)
selfish_uncle_margin.py (does the uncle cap need margin under a private chain?)
spec_point.py (the deployed operating point; the three f-precision arms)
spec_jitter.py (per-recipient delay variance — the transport diagnostic)
Modelling the deployed chain rather than the mechanism
Two defaults are deliberately not spec-faithful, because the report's job is to isolate mechanisms. Flip both for any run meant to answer "what would the deployed chain read":
| knob | default (design) | spec-faithful | why it matters |
|---|---|---|---|
fixed_point / f_precision |
False / 1e6 — exact f |
True / 1000 |
the spec quantises the target rate at 1e3, which reads +1.0 % high; this is the largest error in the deployed estimator |
uncle_model |
countable — the spec's rules |
(same) | --old is an unreachable ceiling, not an alternative: the spec rejects a block carrying a reference that fails the counting rules |
scripts/spec_point.py runs both arms side by side.