9 Commits

Author SHA1 Message Date
Marcin Pawlowski
8cc682aa49
Item 16: the revenue-optimal adversary is not the estimator's worst case
Both eta ceilings in sec 6.6 come from adversaries optimising something else
(revenue, reorg depth), so they bound eta from above without bounding the
damage from below. Optimising the estimate directly needs no ratio transform:
each transition consumes exactly one block-finding event, so minimising
D-hat = (canonical + p_ref * countable uncles)/events is a plain average-reward
MDP over the transition table that already carries the orphan counts. One
value-iteration pass, no bisection.

Unconstrained, the answer degenerates -- and usefully. The optimum is pure
abstention: publish nothing, adopt when overtaken, D-hat = 1 - alpha exactly,
revenue zero. That is sec 6.4's withholding, which the report already shows is
CORRECT measurement rather than mis-measurement, so the unconstrained objective
asks the wrong question.

The constrained one bites. Sweeping lam * (adversary blocks) - (contribution to
D-hat) enumerates policies; the line of interest is where revenue SHARE reaches
alpha, i.e. where attacking costs nothing versus mining honestly. At alpha=0.4
such a policy drives D-hat to 0.642 where the revenue-maximiser reaches 0.811
-- 17 points of extra deflation bought with the selfish premium alone. At 0.36
and 0.45 the gaps are 0.082 and 0.103. Below the 1/3 threshold nothing
profitable deflates, so the exposure starts exactly where selfish mining does.

This revises two claims that were about revenue but read as though they were
about the adversary in general: sec 6.7's "the adversary frontier is exactly
optimal selfish mining; no compounding lever remains" and sec 8.2's echo of it.
Both now say the PROFIT frontier is bounded and the estimator frontier is not
the same policy. Note the sweep parameter is deliberately non-monotone in
revenue -- selfish mining takes a bigger share of a smaller pie, so raw block
rate is maximised by honesty and large lam returns there; it enumerates
policies rather than tracing a path.

Also closes a fairness loop these findings opened. Sec 6.7(1) credits uncle
rewards with compensating orphaned honest producers, computed on the SM1 race
where every orphan is a first-fork block. Under a private chain 20-40% of the
honest blocks destroyed are unreferenceable by construction, so those producers
are uncompensatable at ANY w_u -- not underpaid because p_ref is low, but
unreachable because no valid block may name them. The fairness guarantee
inherits the same first-fork ceiling as the density repair. Logged as item 19,
flagged as a protocol-design question rather than something a schedule fixes.

_solve_mdp is refactored into _solve_reward/_greedy_policy/_stationary/
_policy_rates so both objectives share one implementation; optimal_policy_stats
reproduces its committed figures exactly (eta 0.4413, D-hat 0.9447/0.8111 at
alpha=0.4). 251 tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:54:52 +02:00
Marcin Pawlowski
3708d03e22
Private-chain (SM1) adversary in the per-node engine
Sec 6.8 recorded that "the per-node engine has no private-chain strategy", which is
why every selfish result came from the global race model with uncle recovery as
a free knob eta -- and why open item 5 (does the uncle cap need margin under
attack-inflated orphaning?) could not be sized: a knob has no queue to overflow.

adversary_strategy="selfish" adds it. The coalition mines one shared private
chain and releases under the classic SM1 rules in (a, h) form: adopt when the
public chain wins, match at equal length, override at a one-block lead, else
wait. Only VISIBILITY is modelled -- the coalition's mining needs no special
case, because a member's fork choice already builds on the private tip whenever
it leads (that tip has the greatest height among blocks the member can see) and
falls back to the public chain exactly when the public chain overtakes, which
is the adopt branch. So the private chain forms, extends and is abandoned
emergently, and the code that had to be written is the arrival matrix.

Design notes worth keeping:
- Private blocks reuse the sentinel `withhold` already had (never-arrives), so
  the existing exclusions from canonical-tip selection apply unchanged; release
  flips it back and gossips DIRECTLY from the producer, bypassing Blend, since
  an adversary has no privacy budget to respect and wants the race won.
- A private chain breaks the windowed horizon's premise (a hidden block is old
  enough to look fully-propagated while no honest node has it, and it becomes
  visible LATER, which the one-way frontier pointer cannot revisit), so selfish
  forces the exact full scan and full matrix.
- Blocks still hidden at epoch end are abandoned and hidden from the coalition
  too, or the canonical-tip search would crown a chain no honest node saw.

Validated against Eyal-Sirer at sub-slot latency: revenue share 0.0356 vs an
exact 0.0356 at alpha = 0.1, and above the closed form at higher alpha by just
the margin the alpha_eff fork-amplification correction predicts (0.498 vs 0.484
at alpha = 0.4, with fork rate 0.38).

Adds p_ref_honest: the reference rate over orphans produced OUTSIDE the
coalition. Under a private-chain attack this diverges sharply from p_ref, and
only the honest one measures the repair the report credits to uncle counting --
an attacker's own discarded blocks are its loss to bear.

test_fork unpacks fork_stats positionally, so its three call sites take the new
fifth value. 247 tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 12:46:57 +02:00
Marcin Pawlowski
ef82ed614d
Review pass: reproduce every number from its data of record, fix what did not
Correctness/completeness review of the report and simulator. Verified against
the committed parquets: the sec 6.6 countable-ceiling table (cap-64 MDP sweep),
sec 6.10 Result 4's depth ceilings, the sec 3.4 uncle-selection table, all
adversary-variant numbers, the rho-boundary row-4 quotes (0.976 at rho=0.91,
4-sigma shortfall at 0.96, max cell 1.0024), and the sec 8.4 capstone table.
Three defects found, all fixed:

1. The collapse event was not reproducible from the committed script. Study D
   swept only the default (random) coalition, but the one observed collapse is
   a whale cell; the "once in 144 runs" count came from an ad-hoc probe. The
   committed sweep now carries the selection axis (96 runs) and reproduces the
   event: 1/12 in the whale 50% cell at delta_max = 8, never at 4. All six
   fold-related passages now quote the committed sweep, which also retires the
   stale "the full dynamics never reach it" wording in the sec 6 arc, the
   sec 6.2 intro, row 6 and item 1 -- text that contradicted item 18 since
   yesterday's finding.

2. capstone.py's printout could not reproduce the report's sec 8.4 table. The
   report's numbers are a per-replicate-tail aggregation (each replicate burns
   in against its own early-stop length); the script cut the tail at the ARM's
   max epoch, silently dropping any replicate that stopped earlier (7 of 8 in
   the adversary arm) and landing one rounding step off on three cells. The
   script now aggregates per replicate and prints the SEM; against the existing
   parquet it reproduces the table exactly (1.001/0.998, 0.342+-0.009 /
   0.343+-0.005, p_ref 1.000/0.990, 8 reps both arms). The report table was
   right all along; sec 6.8's p_ref quote (0.989, the per-arm value) is aligned
   to 0.990.

3. Small report fixes: slow-beta deflation rounded 0.765 -> "0.77" (now 0.76);
   fig13's caption now points at the fig36 ceiling instead of implying free
   recovery; row 5 cites the measured slow-beta standing deflation; the
   canonical-data paragraph lists the new studies' artifacts; the simulator
   README's layout block lists the new tests and scripts.

Adds a unit test for reorg.countable_recovery_from_depths (the one new
function that had none). 236 tests pass; the new-study parquets are copied to
the main checkout's runs/, where every other study's data of record lives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 12:29:13 +02:00
Marcin Pawlowski
88f31340ad
Uncle selection: the spec fixes oldest-first, so measure deviation from it
Open item 11 listed "a random (rather than oldest-first) uncle-selection draw"
as an untested spec sensitivity. The spec does not leave it open: Uncle
Selection in cryptarchia-v1-protocol.md has the proposer take the oldest
candidates first, deterministically, because an uncle expires w_u slots after
its own slot. That is exactly what every result in the report already uses, so
the item is a conformance match, not a gap -- and the simulator comment calling
uncle_random_p "the spec's unbiased coin" cites text the spec no longer has.

What is genuinely open is deviation FROM that rule: selection is proposer-local
and the uncles field is never validated. configs/uncle-selection.yaml measures
the cost. A proposer that includes each candidate on a fair coin instead loses
up to 0.10 in D-hat/D, and 0.063 at the recommended W = 10 once rho ~ 1
(0.902 vs 0.965, t = -8.6). At the design point the margin survives but is
spent: 0.980 vs 0.997 against a 0.98 bar. The loss does not close as W grows,
because a coin wastes opportunities rather than queue capacity and a well-sized
window is precisely what keeps the queue short enough for that to bite.

This matters for the sec 8.5 reward recommendation: the spec argues a proposer
has no incentive to deviate BECAUSE uncles grant no reward, and paying them
removes that argument.

Also adds adversary_selection=whale (the largest holders at matched stake, for
the untested concentration case). The marker is appended to key() only when
non-default so every historical run's seed stays byte-identical, guarded by a
test alongside the paired_streams one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:42:16 +02:00
Marcin Pawlowski
0510b20684
Countable recovery under a selfish adversary: SM1 hides the first-fork cost
The countable model can reference only the first block of a fork, so a
discarded chain of h honest blocks yields one countable uncle, not h. Sec 6.6
reads the estimator repair off a free knob eta and quotes it at eta = 1 --
attainable under SM1, which acts the moment the honest branch reaches length 1
and so never buries a second block. The optimal SSZ policy waits and does bury
them, and there the deployed counting rules cap eta at 0.44 (alpha = 0.4,
gamma = 0), landing D-hat at 0.81 rather than the 0.94 an unrestricted count
gives -- and the ceiling degrades with alpha while the unrestricted value
improves. So SM1 is a faithful proxy for selfish-mining revenue (0.484 vs
0.489) but not for TSI's estimator damage.

selfish_mdp: carry per-branch orphan counts on the transition table so the
accounting cannot drift from the race logic; optimal_policy_stats solves the
policy's stationary distribution for per-event canonical/orphan rates. The
per-event rates sum to 1 (every block is canonical or orphaned), which the
tests assert as an independent check on the whole derivation.

reorg: the same ceiling for the depth-maximising adversary -- 0.52 at
alpha = 0.30 with the measured honest fork rate -- reached from the other
direction. Neither adversary optimises deflation directly, so both ceilings
are upper bounds on eta; that gap is logged as open item 16.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:29:57 +02:00
Marcin Pawlowski
4baccd8d4b
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.

Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.

The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.

Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
  taken, so shared streams must give bit-identical trajectories. All 200
  replicate pairs differ by exactly 0.0. Unpaired, the same control only
  had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
  +-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
  chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
  Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
  delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
  both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
  pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
  the same quantity (t = 2.8) had failed correction.

So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.

Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
  reached the parquet; plot_fine_delay.py falls back to the unpaired
  test when it cannot confirm pairing, so the sweep would have completed
  and quietly reported the old result. Caught before the run finished;
  the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
  gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
  itself when the streams are shared, and falls back to the t-test only
  when there is real spread.

§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.

Tests: 214 passed (was 209). ruff clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
Marcin Pawlowski
ac6a309e58
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.

Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
  0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
  differences. They are not resolvable: t = 0.46 and 0.47 over 5
  replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
  Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
  0.996), so the claim is dropped. The 32s curve is non-monotonic with
  replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
  combined first-fork and capacity loss, which this measurement cannot
  separate. Hedged to match fig32's own axis label.

Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.

New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.

Code:
- deep_ref_share is identically 0 on every real countable run: for a
  chain block B the producer's chain below B is the counting chain
  below B, so the counting-side parent-on-chain re-check cannot reject
  what selection emitted. It is a drift alarm, not a rate. Documented
  as such in measure.py, the plot docstring and the config header, and
  pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
  production never called, while carrying most of the selection test
  coverage. Tests now drive select_uncles_at_production through an
  annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
  selection: the pmin/below chain walk that resolves parent-on-chain
  for candidates whose parent sits below the window, and the
  occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
  untested. Now used (the prediction figure reconstructs q_u through
  the identity the report quotes) and tested. The window_miss_prob test
  records that its "~ e^-W" docstring is the f->0 limit: the true decay
  is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
  fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.

Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.

Tests: 209 passed (was 202).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
Marcin Pawlowski
bd2ac7b7be
Countable uncle model: spec counting rules, sweeps, figures
Implement the countable uncle model from the Cryptarchia spec's
counting-only reference rules, and make it the simulator default.

Counting rules (uncles.py, measure.py):
- Only the first block of a fork (parent on the producer's chain) is
  referenceable and countable, which makes every reference verifiable
  from chain data alone.
- The reference window is derived from a window-absorption parameter,
  w_u = W_abs/f slots (W_abs in expected block-intervals, default 10,
  bounded W_abs <= 0.6*k), replacing the free-standing uncle_window.
- Selection skips slots already occupied on the producer's chain and
  takes at most one uncle per slot.
- The measurement pass re-checks every rule per reference and tallies
  rejections as deep_ref_share.

The pre-redesign model is preserved behind --old on tsi-sweep and
tsi-verify. Its RNG key is byte-identical to the pre-uncle_model key,
so --old bit-reproduces the historical runs.

Supporting changes: uncle_model and window_absorption config surface
with validation (config.py, constants.py); accuracy closed form over
the effective q_u (theory.py); plumbing through tsi.py, epoch.py,
sweep.py, blocktree.py, metrics.py, verify.py, figures_pernode.py.

Studies and figures:
- configs/countable-vs-old.yaml -- delay x U grid, run under both
  models on the same grid.
- configs/absorption-window.yaml -- accuracy vs W_abs at U=1.
- scripts/plot_countable_vs_old.py renders fig30-fig33 into
  reports/tsi/report-figures/.

Tests: tests/test_countable_counting.py (7 cases) covering first-fork
eligibility, derived-window bounds, occupied-slot exclusion, and
per-reference re-checking; extensions to test_uncles.py,
test_config.py, test_slot_counting.py. Full fast suite: 202 passed.

Also adds CLAUDE.md (graphify project instructions) and ignores
editor/local-agent state plus the vendored Equi-X benchmark clone.

The reports/tsi/ prose describing this model is held back for a
separate editorial pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 18:48:46 +02:00
Marcin Pawlowski
24da2fc8b3
Importing tsi-sim v3 2026-07-30 18:57:10 +02:00