11 Commits

Author SHA1 Message Date
Marcin Pawlowski
feafe8ec92
Close the sec 6.5 scope variants; static withholding can reach the fold
Whale coalitions, jitter > 0 and very slow beta were the residual "untested
adversary variants" of open item 11. None moves a conclusion:

- Concentration does not change the deflation (suppression D-hat within noise
  at every stake), and a whale coalition reproduces the sec 6.4 withholding law
  D-hat -> (1-beta_adv) more cleanly than a random one: 0.9005/0.6997/0.5010
  against a predicted 0.9/0.7/0.5. The "lumpier share statistic" worry points
  the other way, and for a reason that is about coalition CONSTRUCTION rather
  than concentration: a random coalition grows until its stake first reaches
  the target, so the last node added overshoots by its own size -- a whale,
  under a Pareto tail. Realised block share at a nominal beta_adv = 0.1 is
  0.137 +- 0.108. Logged as item 17: the beta_adv axis is a nominal target.
- jitter up to 1 slot changes nothing under attack (notch 0.390 -> 0.410,
  attacker share flat, range_ratio identically 0), as sec 6.1 found honestly.
- Slow beta shrinks the notch (0.415 -> 0.080 for beta 1 -> 0.1) at flat
  attacker take, but sinks the MEAN estimate to 0.765 at beta = 0.05: the
  estimator can no longer track back up during the honest half of the cycle.
  Slowing beta buys the defender nothing on either axis.

Unplanned: study A blew past the memory guard, which turned out to be the
sec 6.2 fold being reached. The mechanism is sec 6.2's own -- rho_eff = rho/r,
and withholding deflates r by design, so a 50 % coalition doubles the load onto
rho_eff ~ 1.1 at the design point. Swept directly, the estimate collapses once
in 144 runs at delta_max = 8 (a concentrated 50 % coalition) and never at
delta_max = 4. That retires "not an observed dynamical trap" but is one event,
so the claim is stated as a rare tail and the rate is logged unmeasured as
item 18.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 16:53:56 +02:00
Marcin Pawlowski
1a1548b8e5
Re-check sec 8.5 against the spec: the header-padding premise is superseded
Implication (i) quoted "the length of the header does not reveal how many
uncles a block references" from Uncle References. That sentence was removed
from the spec in b809df59 ("Making uncles variable size instead of fixed"):
the field is now a variable-size unpadded list and proposal indistinguishability
is preserved at the message layer, by padding every dispersed payload to
Max_Body_Length. The constraint on a per-reference nephew reward therefore
rests on the voucher being content-dependent -- which the Anonymous Leaders
Reward Protocol still requires it not to be -- and not on header length.

Implications (ii) and (iii) were re-checked against the same revision and stand
unchanged: the "no incentive to deviate" sentence is still there, and the
equal-share content-independent voucher is still the payout model.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:43:13 +02:00
Marcin Pawlowski
88f31340ad
Uncle selection: the spec fixes oldest-first, so measure deviation from it
Open item 11 listed "a random (rather than oldest-first) uncle-selection draw"
as an untested spec sensitivity. The spec does not leave it open: Uncle
Selection in cryptarchia-v1-protocol.md has the proposer take the oldest
candidates first, deterministically, because an uncle expires w_u slots after
its own slot. That is exactly what every result in the report already uses, so
the item is a conformance match, not a gap -- and the simulator comment calling
uncle_random_p "the spec's unbiased coin" cites text the spec no longer has.

What is genuinely open is deviation FROM that rule: selection is proposer-local
and the uncles field is never validated. configs/uncle-selection.yaml measures
the cost. A proposer that includes each candidate on a fair coin instead loses
up to 0.10 in D-hat/D, and 0.063 at the recommended W = 10 once rho ~ 1
(0.902 vs 0.965, t = -8.6). At the design point the margin survives but is
spent: 0.980 vs 0.997 against a 0.98 bar. The loss does not close as W grows,
because a coin wastes opportunities rather than queue capacity and a well-sized
window is precisely what keeps the queue short enough for that to bite.

This matters for the sec 8.5 reward recommendation: the spec argues a proposer
has no incentive to deviate BECAUSE uncles grant no reward, and paying them
removes that argument.

Also adds adversary_selection=whale (the largest holders at matched stake, for
the untested concentration case). The marker is appended to key() only when
non-default so every historical run's seed stays byte-identical, guarded by a
test alongside the paired_streams one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:42:16 +02:00
Marcin Pawlowski
0510b20684
Countable recovery under a selfish adversary: SM1 hides the first-fork cost
The countable model can reference only the first block of a fork, so a
discarded chain of h honest blocks yields one countable uncle, not h. Sec 6.6
reads the estimator repair off a free knob eta and quotes it at eta = 1 --
attainable under SM1, which acts the moment the honest branch reaches length 1
and so never buries a second block. The optimal SSZ policy waits and does bury
them, and there the deployed counting rules cap eta at 0.44 (alpha = 0.4,
gamma = 0), landing D-hat at 0.81 rather than the 0.94 an unrestricted count
gives -- and the ceiling degrades with alpha while the unrestricted value
improves. So SM1 is a faithful proxy for selfish-mining revenue (0.484 vs
0.489) but not for TSI's estimator damage.

selfish_mdp: carry per-branch orphan counts on the transition table so the
accounting cannot drift from the race logic; optimal_policy_stats solves the
policy's stationary distribution for per-event canonical/orphan rates. The
per-event rates sum to 1 (every block is canonical or orphaned), which the
tests assert as an independent check on the whole derivation.

reorg: the same ceiling for the depth-maximising adversary -- 0.52 at
alpha = 0.30 with the measured honest fork rate -- reached from the other
direction. Neither adversary optimises deflation directly, so both ceilings
are upper bounds on eta; that gap is logged as open item 16.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:29:57 +02:00
Marcin Pawlowski
06364797a8
Pair the overload grid; the first-fork cost is resolved at every load
§3.2 rested on 5 unpaired replicates while §3.2a used a paired design.
Re-running the delta_max 4/8/16/32 grid with common random numbers and
20 replicates (configs/countable-vs-old-paired.yaml) changes the answer
at the design end.

The cost is resolved at EVERY delay and grows monotonically with load:

  delta_max=4  (rho 0.36)  -0.0013  t= 4.0   (unpaired: not resolved)
  delta_max=8  (rho 0.56)  -0.0034  t= 9.8   (unpaired: not resolved)
  delta_max=16 (rho 0.96)  -0.0102  t=17.5
  delta_max=32 (rho 1.76)  -0.0228  t=22.4

So §3.2's claim that "at the operating loads (rho < 1) no difference
between the models is detectable at all" was an artefact of the weak
design, not a property of the system. There is a difference; it is just
small — 0.13% and 0.34% at the two sub-unit loads. The new delta_max=4
figure (-0.0013 at 20 reps) independently reproduces §3.2a's (-0.0011 at
40 reps) from a separate sweep.

11 of 12 U>=1 cells resolve individually; max t = 29.0 against a
Bonferroni threshold of 2.87 for twelve tests. The U=0 control is exact:
80/80 replicate pairs differ by 0.0.

New finding at U=1 under overload: the sign FLIPS and the countable rule
wins, +0.0127 (t = 7.6), positive in 19 of 20 pairs. Both models have
collapsed at rho ~ 1.76 with a single uncle slot, but when capacity is
the binding constraint the countable rule's occupied-slot exclusion
means its one reference always recovers a NEW slot, while the
unrestricted rule dedups by block id and can spend that reference on an
orphan whose slot is already counted. Measured recovery agrees:
q_u = 0.591 countable vs 0.579 unrestricted. The slot-vs-block
distinction of §2.1 is worth most exactly where references are scarcest.

Code: paired_gaps and pooled_by_delay move from scripts/plot_fine_delay
into figures_pernode so both plot scripts share one implementation;
plot_countable_vs_old now detects paired runs and uses the
per-replicate difference, falling back to the unpaired two-sample test
otherwise. Two hardcoded reporting values fixed — the Bonferroni
threshold was pinned to 2.935 and printed nan for any grid that was not
15 cells, and a per-cell comparison line had a hardcoded /15 denominator;
both now derive from the grid actually run.

Figures 30-32 regenerated from the paired grid. Tests: 214 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 13:09:21 +02:00
Marcin Pawlowski
3f6f2c4e79
Contents: plain nested list, no HTML entities
The previous version used &nbsp;&nbsp; to fake the indent of subsection
lines. That is an HTML entity leaking into markdown: Obsidian renders it
literally rather than as leading space, so the block reads as broken
there.

Rebuilt as a genuine nested markdown list, one subsection per line,
indented four spaces. Four rather than two because python-markdown (used
by scripts/build_html.py) needs four for reliable nesting, while
Obsidian and GitHub accept it equally — so the same source renders
correctly in all three.

Verified the nesting is real in the built HTML (nested <ul> inside <li>,
not a flattened list) and that the block still covers all 47 anchors
with none dangling and none missing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 12:14:35 +02:00
Marcin Pawlowski
c8a23fd1df
Report becomes the directory README; rebuild the contents block
The report is now reports/tsi/README.md, so browsing to reports/tsi/
lands on the report itself rather than on an index that points at it.
The old index carried nothing the report lacked except the note that
report-figures/ holds the figures of record (the simulator folder does
not commit its own), which is folded into §9; its section table is
superseded by the report's own contents block. The "[Index]" self-link
in the header is replaced by the simulator link the index used to carry.

The contents block was inconsistent: §2 listed subsection titles, §3 and
§6 listed bare numbers with no titles at all, and the appendices were
crammed onto one line while their subsections went unlisted. Rebuilt
from the document's actual headings so every entry has a real title,
top-level entries carry a one-line gloss, and subsections sit indented
under their parent. It now covers all 47 anchors, including B.1-B.4 and
C.1-C.2 which were previously absent.

scripts/build_html.py follows the rename (DOCS is a single document) and
still renders clean: 47 anchors, 0 broken internal links, 0 unrewritten
.md links, 37 images.

Also adds configs/countable-vs-old-paired.yaml — the paired, 20-replicate
version of the overload grid. §3.2a now rests on a paired design while
§3.2 still rests on 5 unpaired replicates, which is why its U=1 cells at
delta_max 16 and 32 sit unresolved at t ~ 0.5 against a replicate sd of
0.15. That sweep is running; the report is not yet updated from it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 12:12:03 +02:00
Marcin Pawlowski
c5059c2fc8
Consolidate the TSI report into one document
The four-part split existed because the single report had grown dense
and heavily cross-referenced; splitting traded that for a different
cost, which the merged read makes visible. Section numbers (§1–§9,
Appendices A–C) were already the stable identifiers, so the parts were
a packaging choice, not a structural one.

reports/tsi/tsi-report.md is now the whole report. Parts are
interleaved back into section order — §1, §2–§5, §6, §7–§8, §9 +
appendices — which is NOT concatenation order: Part 1 carried §1, §7
and §8, so appending files in sequence would have put §7–§8 ahead of
§2. Every cross-file link collapses to an internal anchor; all 47
anchors resolve, all 37 figure embeds resolve, and no line of prose was
lost (verified by diffing normalised content lines with link targets
stripped — 0 lost, additions are the new header and table of contents).

Coherence fixes the merge exposed, all artefacts of the split:
- The roadmap paragraph described "four parts (see the index)" and is
  now a section-order roadmap, with its circular self-link to §1
  dropped.
- §7's figure-location note pointed readers at "the other parts". It
  now names the actual sections, and it was also WRONG about three
  figures: fig17/fig18/fig21 are in Appendix C and figB1/figB2 in
  Appendix B, not §9. It had also never been updated for fig30–fig35.
- §9's "throughout this part" is now "throughout".

README.md becomes a proper index — a section table pointing into the
one document — rather than a list of four files.

scripts/split_report.py is deleted: a one-time migration that produced
the split, now both obsolete and pointing the wrong way.

scripts/build_html.py was already broken before this change — it still
read the report from tsi-sim-pernode/, where the files stopped living
when they moved to reports/tsi/. Retargeted at reports/tsi/ and the
single document; verified end-to-end (0 broken internal anchors, 0
unrewritten .md links, 37 images in the rendered HTML). Its output is
now gitignored, as its docstring always claimed it was.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 10:42:03 +02:00
Marcin Pawlowski
ac6a309e58
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.

Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
  0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
  differences. They are not resolvable: t = 0.46 and 0.47 over 5
  replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
  Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
  0.996), so the claim is dropped. The 32s curve is non-monotonic with
  replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
  combined first-fork and capacity loss, which this measurement cannot
  separate. Hedged to match fig32's own axis label.

Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.

New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.

Code:
- deep_ref_share is identically 0 on every real countable run: for a
  chain block B the producer's chain below B is the counting chain
  below B, so the counting-side parent-on-chain re-check cannot reject
  what selection emitted. It is a drift alarm, not a rate. Documented
  as such in measure.py, the plot docstring and the config header, and
  pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
  production never called, while carrying most of the selection test
  coverage. Tests now drive select_uncles_at_production through an
  annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
  selection: the pmin/below chain walk that resolves parent-on-chain
  for candidates whose parent sits below the window, and the
  occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
  untested. Now used (the prediction figure reconstructs q_u through
  the identity the report quotes) and tested. The window_miss_prob test
  records that its "~ e^-W" docstring is the f->0 limit: the true decay
  is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
  fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.

Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.

Tests: 209 passed (was 202).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
Marcin Pawlowski
cb58cd7ead
Round-4 TSI report review: apply findings, editorial pass, code + figure fixes
Applied the reconstructed round-4 review to the TSI parameter-selection report
set (reports/tsi) and executed the follow-ups.

Report (reports/tsi):
- Applied the must+should findings across README + parts 1-4: cross-part numeric
  corrections, figure-caption fixes, spec reconciliation, and cross-file companions
  (hops-degradation and notch/reward numbers, tip-agreement ordering, density-window
  timing, VRF -> ZK Proof-of-Leadership, w_u window/reward gloss).
- Editorial pass for timeless voice (no "now adopted / merged / coin" narration) and
  a gentle spec-safety framing (recommendations are thresholds; the protocol's
  MAX_UNCLES=4 sits safely above them).
- Added the fork-rate-vs-scale table (6.10), defined "grinding gain", promoted the
  clock-skew study to its own paragraph, added the correlated-latency caveat, and
  moved fig27/fig28 beside their discussion.
- Documented the Blend cascade in 2: hops propagate over the shared gossip graph
  (not direct links), the final broadcast comes from the last relay, relays are
  blind forwarders.

Simulator (tools/simulators/tsi/tsi-sim-pernode):
- Docstring/dead-code fixes: theory.block_count_ceiling (legacy framing), measure,
  reorg (catch-up reading), metrics (removed two dead helpers), config (fixed_point
  10^-6; clock_skew_max/lottery_chunks documented inert), stake_vs_delay.
- Generator correctness + regenerated figures: figures_pernode.CONFIG_COLS now
  exhaustive (f no longer pooled); rho_boundary_analysis SEM across replicates +
  hollow floored markers + de-hardcoded ell_mean (measured from the run's graph);
  appendix_fluct per-N sigma + ~18x title (figB2); bootstrap_dynamics driving
  estimate so fig1 epoch-0 matches genesis.
- pytest: 186 passed; report links 528/0 dangling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-31 13:13:03 +02:00
Marcin Pawlowski
43d09b8fa6
Importing TSI report 2026-07-30 18:23:59 +02:00