Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
"""High-precision figures for the LOW mixing-delay band (the design regime).
|
|
|
|
|
|
|
|
|
|
|
|
countable-vs-old.yaml samples delay at 4/8/16/32 with 5 replicates. That resolves the
|
|
|
|
|
|
overload regime but leaves the design regime under-measured: every countable-vs-unrestricted
|
|
|
|
|
|
gap at delay <= 8 sits inside the replicate noise there, so the only honest statement is
|
|
|
|
|
|
"no difference detected" — with no bound on how large an undetected difference could be.
|
|
|
|
|
|
|
|
|
|
|
|
fine-delay.yaml spends replicates instead of range (delay 1..5, 40 replicates) to turn that
|
|
|
|
|
|
into a real bound. Consumes:
|
|
|
|
|
|
|
|
|
|
|
|
tsi-sweep --config configs/fine-delay.yaml --label fine-countable
|
|
|
|
|
|
tsi-sweep --config configs/fine-delay.yaml --old --label fine-old
|
|
|
|
|
|
|
|
|
|
|
|
and renders (into --out):
|
|
|
|
|
|
|
|
|
|
|
|
fine_accuracy_vs_delay equilibrium D/D_true vs delay 1..5, countable (solid) vs
|
|
|
|
|
|
unrestricted (dashed) per U, replicate-SEM bars.
|
|
|
|
|
|
fine_gap_vs_delay THE precision figure: the countable - unrestricted gap with 95%
|
|
|
|
|
|
CIs, against the U=0 negative-control band. A CI straddling zero
|
|
|
|
|
|
means no difference at this power; the band shows the floor.
|
|
|
|
|
|
|
|
|
|
|
|
Usage:
|
|
|
|
|
|
python scripts/plot_fine_delay.py --countable RUNDIR --old RUNDIR \
|
|
|
|
|
|
--out figures/fine-delay
|
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
|
|
import argparse
|
|
|
|
|
|
from pathlib import Path
|
2026-08-05 13:09:21 +02:00
|
|
|
|
from statistics import NormalDist
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
|
|
|
|
|
|
import matplotlib.pyplot as plt
|
|
|
|
|
|
import numpy as np
|
|
|
|
|
|
import pandas as pd
|
|
|
|
|
|
|
|
|
|
|
|
from tsi_sim.plotting import style
|
2026-08-05 13:09:21 +02:00
|
|
|
|
from tsi_sim.plotting.figures_pernode import (
|
|
|
|
|
|
Z95,
|
|
|
|
|
|
equilibrium,
|
|
|
|
|
|
paired_gaps,
|
|
|
|
|
|
pooled_by_delay,
|
|
|
|
|
|
rho_for,
|
|
|
|
|
|
sem,
|
|
|
|
|
|
)
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
|
|
|
|
|
|
DELAY = "blend_delay_max"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _load(run_dir: str | Path) -> pd.DataFrame:
|
|
|
|
|
|
return pd.read_parquet(Path(run_dir) / "results.parquet")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _cells(df: pd.DataFrame) -> pd.DataFrame:
|
|
|
|
|
|
"""Per (delay, U): replicate mean, SEM and count of the equilibrium accuracy."""
|
|
|
|
|
|
return equilibrium(df).groupby([DELAY, "max_uncles"], as_index=False).agg(
|
|
|
|
|
|
mean_ratio=("mean_ratio", "mean"),
|
|
|
|
|
|
sem_ratio=("mean_ratio", sem),
|
|
|
|
|
|
n_rep=("mean_ratio", "size"),
|
|
|
|
|
|
mean_q=("mean_q", "mean"),
|
|
|
|
|
|
mean_q_eff=("mean_q_eff", "mean"))
|
|
|
|
|
|
|
|
|
|
|
|
|
Correctness pass: derive rho in code, and the absolute vs-1.0 test
Two findings from re-reviewing the fine-delay section.
1. The rho values I put in s3.2a were wrong. The report derives
rho = f*D_vis with D_vis = hops*delta_max/2 + (hops+1)*ell_mean from
a MEASURED ell_mean (1.211 slots at N=1000/degree=6), not from the
link_latency_mean parameter (0.5). Hand-substituting a guessed 1.5
inflated every value by ~0.04: the band is rho 0.21-0.41, not
0.25-0.45.
To stop that recurring, graph_ell_mean moves out of
rho_boundary_analysis.py into figures_pernode.py, joined by a new
rho_for() that both scripts and any future quotation go through;
plot_fine_delay.py now prints the derived rho per delay.
This also exposed an inconsistency in the existing s3.2 table, which
rounded delta_max=4 to "rho ~ 0.4" while s3.2a called the same cell
0.36 and prose elsewhere already used 0.56 for delta_max=8. The s3.2
column now carries the derived values (0.36/0.56/0.96/1.76).
2. Testing each cell against the exact target 1.0 -- the same question
the gap test asks, without reference to the other model --
corroborates the first-fork onset independently. Unrestricted: 1/15
cells below 1 (t=-2.09, chance). Countable: 4/15, and not scattered
-- delta_max=4 at U=1, and ALL THREE caps at delta_max=5 (-0.0012 to
-0.0019, t=-2.5..-3.7). A shortfall appearing at every cap at once,
only at the top of the band, only under the restricted model, is the
first-fork cost seen absolutely.
That makes "one uncle slot is sufficient -- not approximately,
exactly" too strong as I had written it. s3.2a now states the
residual (0.1-0.2% at the top of the band, zero below delta_max=3),
reconciles it with the s1 headline, and notes that since all three
caps show the same shortfall the residual is not a capacity limit.
The bound quoted in s1 moves from "below 0.15%" to "<= 0.2%".
Also adds the new run directories to s9's canonical list, which covered
every other study but not these.
Tests: 209 passed. ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 10:34:33 +02:00
|
|
|
|
def vs_one(cells: pd.DataFrame) -> pd.DataFrame:
|
|
|
|
|
|
"""Test each U >= 1 cell against the exact target 1.0.
|
|
|
|
|
|
|
|
|
|
|
|
An independent read on the same question the gap test asks: the report's claim is that
|
|
|
|
|
|
uncle recovery restores the equilibrium to EXACTLY the true stake, so a systematic
|
|
|
|
|
|
shortfall across uncle caps is the first-fork cost seen from the absolute side rather
|
|
|
|
|
|
than differentially. 40 replicates give ~0.0005 resolution, enough to see 0.1%.
|
|
|
|
|
|
"""
|
|
|
|
|
|
u = cells[cells.max_uncles > 0].copy()
|
|
|
|
|
|
u["dev"] = u.mean_ratio - 1.0
|
|
|
|
|
|
u["t"] = u.dev / u.sem_ratio.replace(0, np.nan)
|
|
|
|
|
|
return u.sort_values(["max_uncles", DELAY])
|
|
|
|
|
|
|
|
|
|
|
|
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
def gaps(cnt: pd.DataFrame, old: pd.DataFrame) -> pd.DataFrame:
|
|
|
|
|
|
"""countable - unrestricted per cell, with the unpaired SE and 95% CI half-width."""
|
|
|
|
|
|
m = cnt.merge(old, on=[DELAY, "max_uncles"], suffixes=("_c", "_o"))
|
|
|
|
|
|
m["gap"] = m.mean_ratio_c - m.mean_ratio_o
|
|
|
|
|
|
m["se"] = np.hypot(m.sem_ratio_c, m.sem_ratio_o)
|
|
|
|
|
|
m["ci95"] = Z95 * m.se
|
|
|
|
|
|
m["t"] = np.where(m.se > 0, np.abs(m.gap) / m.se.replace(0, np.nan), np.inf)
|
|
|
|
|
|
return m.sort_values(["max_uncles", DELAY])
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def fig_accuracy(cnt: pd.DataFrame, old: pd.DataFrame) -> plt.Figure:
|
|
|
|
|
|
fig, ax = plt.subplots(figsize=(6.4, 4.2))
|
|
|
|
|
|
for i, u in enumerate(sorted(cnt["max_uncles"].unique())):
|
|
|
|
|
|
c = style.color_for(i)
|
|
|
|
|
|
a = cnt[cnt.max_uncles == u].sort_values(DELAY)
|
|
|
|
|
|
b = old[old.max_uncles == u].sort_values(DELAY)
|
|
|
|
|
|
ctl = " (control)" if u == 0 else ""
|
|
|
|
|
|
ax.errorbar(a[DELAY], a.mean_ratio, yerr=a.sem_ratio, fmt="-o", color=c,
|
|
|
|
|
|
label=f"U={u} countable{ctl}", ms=4, capsize=2, lw=1.2)
|
|
|
|
|
|
ax.errorbar(b[DELAY], b.mean_ratio, yerr=b.sem_ratio, fmt="--s", color=c,
|
|
|
|
|
|
label=f"U={u} unrestricted{ctl}", ms=4, capsize=2, alpha=0.75, lw=1.2)
|
|
|
|
|
|
ax.axhline(1.0, color="0.4", lw=0.8, ls=":")
|
|
|
|
|
|
ax.set_xlabel("max per-relay mixing delay (slots)")
|
|
|
|
|
|
ax.set_ylabel(r"equilibrium $\hat{D}/D_{true}$")
|
|
|
|
|
|
ax.set_title("Design-regime accuracy: countable vs unrestricted referencing")
|
|
|
|
|
|
ax.legend(ncol=2, fontsize="x-small")
|
|
|
|
|
|
return fig
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def fig_gap(g: pd.DataFrame) -> plt.Figure:
|
|
|
|
|
|
"""The countable − unrestricted gap with 95% CIs, zoomed to the U >= 1 scale.
|
|
|
|
|
|
|
|
|
|
|
|
The U=0 negative control is NOT plotted as a band here: with no uncles the models are
|
|
|
|
|
|
identical by construction, but the unrecovered regime is so noisy that its CI (±0.025)
|
|
|
|
|
|
is ~17x the entire range of the U >= 1 gaps and would fill the axes. Its magnitude is
|
|
|
|
|
|
annotated instead — the point being that the control's noise floor lives far outside
|
|
|
|
|
|
anything the uncle arms show, so those arms are measuring signal, not spread.
|
|
|
|
|
|
"""
|
|
|
|
|
|
fig, ax = plt.subplots(figsize=(6.6, 4.2))
|
|
|
|
|
|
pooled = pooled_by_delay(g)
|
|
|
|
|
|
for i, u in enumerate(sorted(g["max_uncles"].unique())):
|
|
|
|
|
|
if u == 0:
|
|
|
|
|
|
continue
|
|
|
|
|
|
a = g[g.max_uncles == u].sort_values(DELAY)
|
|
|
|
|
|
ax.errorbar(a[DELAY], a.gap, yerr=a.ci95, fmt="-o", color=style.color_for(i),
|
|
|
|
|
|
label=f"U={u}", ms=4, capsize=3, lw=1.0, alpha=0.75)
|
|
|
|
|
|
ax.errorbar(pooled[DELAY], pooled.gap, yerr=pooled.ci95, fmt="-D", color="0.15",
|
|
|
|
|
|
label="pooled over U≥1", ms=5, capsize=4, lw=1.8, zorder=5)
|
|
|
|
|
|
ax.axhline(0.0, color="0.3", lw=0.9, ls=":")
|
|
|
|
|
|
ax.set_xlabel("max per-relay mixing delay (slots)")
|
|
|
|
|
|
ax.set_ylabel(r"$\hat{D}/D$ gap: countable $-$ unrestricted")
|
|
|
|
|
|
ax.set_title("The first-fork cost across the design band (95% CI)")
|
|
|
|
|
|
ctl = g[g.max_uncles == 0]
|
|
|
|
|
|
if len(ctl):
|
|
|
|
|
|
band = float(ctl.ci95.max())
|
|
|
|
|
|
span = float(np.abs(np.r_[g[g.max_uncles > 0].gap + g[g.max_uncles > 0].ci95,
|
|
|
|
|
|
g[g.max_uncles > 0].gap - g[g.max_uncles > 0].ci95]).max())
|
|
|
|
|
|
ax.set_ylim(-1.35 * span, 1.35 * span)
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
# Under pairing the control is exactly 0 in every replicate pair (shared streams), so
|
|
|
|
|
|
# quoting a CI for it is meaningless; unpaired, the width of that CI is the point.
|
|
|
|
|
|
note = ("U=0 negative control: exactly 0.0 in all 200 replicate pairs "
|
|
|
|
|
|
"(shared streams — an identity check, not a noise check)"
|
|
|
|
|
|
if band == 0.0 else
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
f"U=0 negative control (true gap = 0): 95% CI ±{band:.4f}, "
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
f"{band / span:.0f}× outside this range")
|
|
|
|
|
|
ax.text(0.015, 0.03, note, transform=ax.transAxes, fontsize=6.5, alpha=0.75)
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
ax.legend(fontsize="x-small", ncol=2)
|
|
|
|
|
|
return fig
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def main() -> None:
|
|
|
|
|
|
ap = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
|
|
|
|
|
ap.add_argument("--countable", required=True, help="run dir of fine-countable")
|
|
|
|
|
|
ap.add_argument("--old", required=True, help="run dir of fine-old")
|
|
|
|
|
|
ap.add_argument("--out", default="figures/fine-delay")
|
|
|
|
|
|
args = ap.parse_args()
|
|
|
|
|
|
style.apply_style()
|
|
|
|
|
|
out = Path(args.out)
|
|
|
|
|
|
out.mkdir(parents=True, exist_ok=True)
|
|
|
|
|
|
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
cnt_raw, old_raw = _load(args.countable), _load(args.old)
|
|
|
|
|
|
cnt, old = _cells(cnt_raw), _cells(old_raw)
|
|
|
|
|
|
unpaired = gaps(cnt, old)
|
|
|
|
|
|
pg = paired_gaps(cnt_raw, old_raw)
|
|
|
|
|
|
# The paired test supersedes the unpaired one when both arms share their streams: same
|
|
|
|
|
|
# estimand, far smaller standard error. Keep the unpaired numbers for the variance-
|
|
|
|
|
|
# reduction report below.
|
|
|
|
|
|
g = unpaired if pg is None else unpaired.drop(columns=["gap", "se", "ci95", "t"]).merge(
|
|
|
|
|
|
pg[[DELAY, "max_uncles", "gap", "se", "ci95", "t"]], on=[DELAY, "max_uncles"])
|
|
|
|
|
|
prov = ("tsi-sim-pernode fine-delay%s.yaml (+--old)"
|
|
|
|
|
|
% ("-paired" if pg is not None else ""))
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
|
Correctness pass: derive rho in code, and the absolute vs-1.0 test
Two findings from re-reviewing the fine-delay section.
1. The rho values I put in s3.2a were wrong. The report derives
rho = f*D_vis with D_vis = hops*delta_max/2 + (hops+1)*ell_mean from
a MEASURED ell_mean (1.211 slots at N=1000/degree=6), not from the
link_latency_mean parameter (0.5). Hand-substituting a guessed 1.5
inflated every value by ~0.04: the band is rho 0.21-0.41, not
0.25-0.45.
To stop that recurring, graph_ell_mean moves out of
rho_boundary_analysis.py into figures_pernode.py, joined by a new
rho_for() that both scripts and any future quotation go through;
plot_fine_delay.py now prints the derived rho per delay.
This also exposed an inconsistency in the existing s3.2 table, which
rounded delta_max=4 to "rho ~ 0.4" while s3.2a called the same cell
0.36 and prose elsewhere already used 0.56 for delta_max=8. The s3.2
column now carries the derived values (0.36/0.56/0.96/1.76).
2. Testing each cell against the exact target 1.0 -- the same question
the gap test asks, without reference to the other model --
corroborates the first-fork onset independently. Unrestricted: 1/15
cells below 1 (t=-2.09, chance). Countable: 4/15, and not scattered
-- delta_max=4 at U=1, and ALL THREE caps at delta_max=5 (-0.0012 to
-0.0019, t=-2.5..-3.7). A shortfall appearing at every cap at once,
only at the top of the band, only under the restricted model, is the
first-fork cost seen absolutely.
That makes "one uncle slot is sufficient -- not approximately,
exactly" too strong as I had written it. s3.2a now states the
residual (0.1-0.2% at the top of the band, zero below delta_max=3),
reconciles it with the s1 headline, and notes that since all three
caps show the same shortfall the residual is not a capacity limit.
The bound quoted in s1 moves from "below 0.15%" to "<= 0.2%".
Also adds the new run directories to s9's canonical list, which covered
every other study but not these.
Tests: 209 passed. ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 10:34:33 +02:00
|
|
|
|
# rho per delay, derived (never hand-substituted) — these are the report's axis labels.
|
|
|
|
|
|
delays = sorted(cnt[DELAY].unique())
|
|
|
|
|
|
print("load rho = f*D_vis per delay (measured ell_mean, see figures_pernode.rho_for):")
|
|
|
|
|
|
print(" " + " ".join(f"delay={d:g}: rho={r:.3f}"
|
|
|
|
|
|
for d, r in zip(delays, rho_for(cnt_raw, delays), strict=True)))
|
|
|
|
|
|
print()
|
|
|
|
|
|
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
written = []
|
|
|
|
|
|
written += style.save(fig_accuracy(cnt, old), out / "fine_accuracy_vs_delay", prov)
|
|
|
|
|
|
written += style.save(fig_gap(g), out / "fine_gap_vs_delay", prov)
|
|
|
|
|
|
|
|
|
|
|
|
print(f"{'cell':<16} {'countable':>17} {'unrestricted':>17} "
|
|
|
|
|
|
f"{'gap':>9} {'95% CI':>9} {'t':>6} verdict")
|
|
|
|
|
|
for _, r in g.iterrows():
|
|
|
|
|
|
verdict = ("CONTROL (true gap = 0)" if r.max_uncles == 0 else
|
|
|
|
|
|
"resolved" if r.t >= 2 else "no difference resolved")
|
|
|
|
|
|
print(f"U={int(r.max_uncles)} delay={r[DELAY]:>5g} "
|
|
|
|
|
|
f"{r.mean_ratio_c:.4f}+-{r.sem_ratio_c:.4f} "
|
|
|
|
|
|
f"{r.mean_ratio_o:.4f}+-{r.sem_ratio_o:.4f} "
|
|
|
|
|
|
f"{r.gap:+.4f} +-{r.ci95:.4f} {r.t:6.2f} {verdict}")
|
|
|
|
|
|
worst = g[g.max_uncles > 0]
|
|
|
|
|
|
print(f"\nn_rep = {int(g.n_rep_c.min())}/{int(g.n_rep_o.min())} per arm")
|
|
|
|
|
|
print(f"widest 95% CI half-width at U>=1: +-{worst.ci95.max():.4f} "
|
|
|
|
|
|
f"({100 * worst.ci95.max():.2f} pp)")
|
2026-08-05 13:09:21 +02:00
|
|
|
|
# Per-cell significance must be read against the number of cells tested — some are
|
|
|
|
|
|
# expected to clear t=2 by chance alone — so quote the Bonferroni threshold for the grid
|
|
|
|
|
|
# actually run, not a constant baked in for one grid size.
|
|
|
|
|
|
bonf = NormalDist().inv_cdf(1 - 0.05 / (2 * len(worst)))
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
print(f"per-cell: {int((worst.t >= 2).sum())}/{len(worst)} cells with t>=2 "
|
|
|
|
|
|
f"(expected by chance {0.05 * len(worst):.2f}); max t = {worst.t.max():.2f} "
|
|
|
|
|
|
f"vs Bonferroni threshold {bonf:.3f}")
|
|
|
|
|
|
|
|
|
|
|
|
print("\npooled over U>=1 (the three caps measure the same difference):")
|
|
|
|
|
|
pooled = pooled_by_delay(g)
|
|
|
|
|
|
for _, r in pooled.iterrows():
|
|
|
|
|
|
mark = " <-- resolved" if r.t >= 2 else ""
|
|
|
|
|
|
print(f" delay={r[DELAY]:>4g}: {r.gap:+.5f} +-{r.ci95:.5f} t={r.t:5.2f}{mark}")
|
|
|
|
|
|
w = 1.0 / worst.se.to_numpy() ** 2
|
|
|
|
|
|
allp = float((worst.gap.to_numpy() * w).sum() / w.sum())
|
|
|
|
|
|
alle = float(np.sqrt(1.0 / w.sum()))
|
|
|
|
|
|
print(f" whole band: {allp:+.5f} +-{Z95 * alle:.5f} t={abs(allp) / alle:.2f}")
|
|
|
|
|
|
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
if pg is not None:
|
|
|
|
|
|
print("\n=== PAIRED design (common random numbers) ===")
|
|
|
|
|
|
# Variance reduction actually achieved, per cell, vs the unpaired standard error.
|
|
|
|
|
|
cmp = unpaired[[DELAY, "max_uncles", "se"]].merge(
|
|
|
|
|
|
pg[[DELAY, "max_uncles", "se", "n_pair"]], on=[DELAY, "max_uncles"],
|
|
|
|
|
|
suffixes=("_unpaired", "_paired"))
|
|
|
|
|
|
u = cmp[cmp.max_uncles > 0]
|
|
|
|
|
|
ratio = (u.se_unpaired / u.se_paired.replace(0, np.nan))
|
|
|
|
|
|
print(f" SE shrink at U>=1: median {ratio.median():.1f}x, range "
|
|
|
|
|
|
f"{ratio.min():.1f}-{ratio.max():.1f}x ({int(u.n_pair.min())} pairs/cell)")
|
|
|
|
|
|
pctl = pg[pg.max_uncles == 0]
|
|
|
|
|
|
exact, tot = int(pctl.n_zero.sum()), int(pctl.n_pair.sum())
|
|
|
|
|
|
print(f" U=0 control under pairing must be EXACTLY zero: {exact}/{tot} pairs are 0.0"
|
|
|
|
|
|
f" -> {'PASSES' if exact == tot else 'FAILS — streams are not shared'}")
|
|
|
|
|
|
res = pg[(pg.max_uncles > 0) & (pg.t >= 2)]
|
2026-08-05 13:09:21 +02:00
|
|
|
|
n_cells = len(pg[pg.max_uncles > 0])
|
|
|
|
|
|
n_unpaired = int((unpaired[unpaired.max_uncles > 0].t >= 2).sum())
|
|
|
|
|
|
print(f" per-cell resolved at |t|>=2: {len(res)}/{n_cells}"
|
|
|
|
|
|
f" (same data, unpaired test: {n_unpaired}/{n_cells})")
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
ctl = g[g.max_uncles == 0]
|
|
|
|
|
|
if len(ctl):
|
Paired design: resolve the design band with common random numbers
The unpaired comparison could not answer the question it was asked. The
two uncle models draw independent RNG streams -- uncle_model is in the
config key, which is what makes --old bit-reproduce earlier runs -- so
the arms differed in stake draw, peering graph and every lottery
outcome, each comparison paid the between-run variance twice, and the
per-cell floor (+-0.0015) sat an order of magnitude above the effect.
Only delta_max = 5 resolved, and only after pooling.
Adds `paired_streams`: the RNG root is derived from the model-
independent part of the key, so a countable cell and its --old twin get
the SAME stake, graph and lottery draws and the uncle rule is the only
difference. Each replicate is then a matched pair and the shared
variance cancels. Trajectories still diverge after epoch 0 through the
genuine feedback (a different counted density changes the next epoch's
difficulty), which is the signal.
The flag is deliberately NOT in key(): it selects which key the seed is
derived from, so including it would perturb every historical seed.
Re-verified that --old still bit-reproduces the committed 2026-07-27
rho-boundary parquet, max |delta| = 0.
Results (configs/fine-delay-paired.yaml, 40 replicates per arm):
- Negative control becomes an IDENTITY check. With U = 0 no reference is
taken, so shared streams must give bit-identical trajectories. All 200
replicate pairs differ by exactly 0.0. Unpaired, the same control only
had to agree within +-0.025 and drifted by 0.016.
- Per-cell SE shrinks by a median 1.6x (1.2-2.1x); widest 95% CI goes
+-0.0015 -> +-0.0010. 5/15 cells resolve at |t| >= 2 (0.75 expected by
chance); the largest, U=2 at delta_max=4, is t = 4.32 and clears
Bonferroni for 15 tests.
- The cost is a STEP, not the ramp the unpaired data suggested:
delta_max 1-3 unresolved (t = 1.1, 1.8, 1.4), then delta_max 4 AND 5
both resolve at -0.0011 (t = 4.7) and -0.0009 (t = 3.7). Whole-band
pooled -0.00060 +- 0.00021, t = 5.7 -- where the unpaired estimate of
the same quantity (t = 2.8) had failed correction.
So the first-fork restriction costs nothing measurable up to
delta_max = 3 and about 0.1% at 4-5 -- an order of magnitude below the
+-0.9% per-epoch sampling noise.
Two bugs found while building this, both of which would have silently
produced a wrong answer:
- paired_streams was missing from metrics._CONFIG_FIELDS, so it never
reached the parquet; plot_fine_delay.py falls back to the unpaired
test when it cannot confirm pairing, so the sweep would have completed
and quietly reported the old result. Caught before the run finished;
the sweep was restarted and a test now pins the field.
- The U=0 control check reported FAILS on a PERFECT control: paired, the
gap is exactly 0 so its SE is 0 and t is 0/0. It now checks the gap
itself when the streams are shared, and falls back to the t-test only
when there is real spread.
§3.2a is rewritten around the paired measurement; the unpaired sweep is
retained in §9 as the power comparison that motivated it. Figures 34-35
regenerated, with the control annotation and provenance reflecting the
design actually used.
Tests: 214 passed (was 209). ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:58:26 +02:00
|
|
|
|
# Under pairing the control gap is EXACTLY 0, so its se is 0 and t is 0/0. That is the
|
|
|
|
|
|
# ideal outcome, not a failure — check the gap itself, and only fall back to the
|
|
|
|
|
|
# t-based check when there is real spread to test (the unpaired case).
|
|
|
|
|
|
worst = float(ctl.gap.abs().max())
|
|
|
|
|
|
exact = bool((ctl.se == 0).all()) if "se" in ctl else False
|
|
|
|
|
|
ok = (worst == 0.0) if exact else (float(ctl.t.max()) < 2)
|
|
|
|
|
|
how = "identical by construction" if exact else "within noise"
|
|
|
|
|
|
print(f"\nU=0 negative control (true gap = 0): |gap| up to {worst:.4g} ({how})"
|
|
|
|
|
|
f" -> control {'PASSES' if ok else 'FAILS'}")
|
Correctness pass: derive rho in code, and the absolute vs-1.0 test
Two findings from re-reviewing the fine-delay section.
1. The rho values I put in s3.2a were wrong. The report derives
rho = f*D_vis with D_vis = hops*delta_max/2 + (hops+1)*ell_mean from
a MEASURED ell_mean (1.211 slots at N=1000/degree=6), not from the
link_latency_mean parameter (0.5). Hand-substituting a guessed 1.5
inflated every value by ~0.04: the band is rho 0.21-0.41, not
0.25-0.45.
To stop that recurring, graph_ell_mean moves out of
rho_boundary_analysis.py into figures_pernode.py, joined by a new
rho_for() that both scripts and any future quotation go through;
plot_fine_delay.py now prints the derived rho per delay.
This also exposed an inconsistency in the existing s3.2 table, which
rounded delta_max=4 to "rho ~ 0.4" while s3.2a called the same cell
0.36 and prose elsewhere already used 0.56 for delta_max=8. The s3.2
column now carries the derived values (0.36/0.56/0.96/1.76).
2. Testing each cell against the exact target 1.0 -- the same question
the gap test asks, without reference to the other model --
corroborates the first-fork onset independently. Unrestricted: 1/15
cells below 1 (t=-2.09, chance). Countable: 4/15, and not scattered
-- delta_max=4 at U=1, and ALL THREE caps at delta_max=5 (-0.0012 to
-0.0019, t=-2.5..-3.7). A shortfall appearing at every cap at once,
only at the top of the band, only under the restricted model, is the
first-fork cost seen absolutely.
That makes "one uncle slot is sufficient -- not approximately,
exactly" too strong as I had written it. s3.2a now states the
residual (0.1-0.2% at the top of the band, zero below delta_max=3),
reconciles it with the s1 headline, and notes that since all three
caps show the same shortfall the residual is not a capacity limit.
The bound quoted in s1 moves from "below 0.15%" to "<= 0.2%".
Also adds the new run directories to s9's canonical list, which covered
every other study but not these.
Tests: 209 passed. ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 10:34:33 +02:00
|
|
|
|
|
|
|
|
|
|
# Absolute test: does uncle recovery actually land on 1.0? Same question as the gap
|
|
|
|
|
|
# test, asked without reference to the other model.
|
|
|
|
|
|
for lbl, cells in (("countable", cnt), ("unrestricted", old)):
|
|
|
|
|
|
v = vs_one(cells)
|
|
|
|
|
|
lo = v[v.t <= -2]
|
|
|
|
|
|
print(f"\nvs exact 1.0, {lbl}: {len(lo)}/{len(v)} cells significantly BELOW 1")
|
|
|
|
|
|
if len(lo):
|
|
|
|
|
|
for d, s in lo.groupby(DELAY):
|
|
|
|
|
|
caps = "/".join(f"U={int(x)}" for x in sorted(s.max_uncles))
|
|
|
|
|
|
print(f" delay={d:>4g}: {caps} dev {s.dev.min():+.5f}..{s.dev.max():+.5f} "
|
|
|
|
|
|
f"t {s.t.min():.2f}..{s.t.max():.2f}")
|
Review fixes + high-precision design-band delay study
Acts on a correctness/completeness review of the countable uncle model
and its report material.
Correctness fixes in the report:
- s3.4 quoted 0.998 for W_abs=10 at the 8s budget; the run says 0.9963.
- s1 claimed both models >= 0.996 at U >= 1; countable U=2 delta=8 is
0.9955. Corrected to >= 0.995.
- The s3.2 table presented two cells (U=1 at delta 16 and 32) as model
differences. They are not resolvable: t = 0.46 and 0.47 over 5
replicates. The table now carries +-SEM and a t per cell.
- s3.4 claimed the ~7-block-interval floor "carries over unchanged".
Accuracy is still climbing past W=7 at every delay (8s: 0.989 ->
0.996), so the claim is dropped. The 32s curve is non-monotonic with
replicate SD up to 0.22 and is now flagged as noise, not a trend.
- 1-r was attributed to the first-fork restriction alone; it is the
combined first-fork and capacity loss, which this measurement cannot
separate. Hedged to match fig32's own axis label.
Completeness: the U=0 negative control was swept but never reported.
With no uncles the two models are identical by construction, yet they
differ by -0.23 at delta_max=32 (t=2.1) because they draw independent
RNG streams. That is the noise floor the rest of the grid must clear,
and it is now in s3.2, s9, fig30 and the config header.
New study (configs/fine-delay.yaml, scripts/plot_fine_delay.py, s3.2a,
fig34/fig35): the design band delta_max 1-5 at 40 replicates, both
models. Findings: every U >= 1 cell of both models lands in
0.998-1.001, flat in delay, while U=0 decays 0.810 -> 0.640. No
individual cell resolves a model difference (widest 95% CI +-0.15pp;
max t=2.59 vs Bonferroni 2.94 over 15 cells). Pooled across uncle caps
the first-fork cost is monotone in delay and separates from zero only
at delta_max=5 (-0.0014 +- 0.0007, t=3.7) -- below 0.15% everywhere in
the band, against +-0.9% per-epoch sampling noise.
Code:
- deep_ref_share is identically 0 on every real countable run: for a
chain block B the producer's chain below B is the counting chain
below B, so the counting-side parent-on-chain re-check cannot reject
what selection emitted. It is a drift alarm, not a rate. Documented
as such in measure.py, the plot docstring and the config header, and
pinned by a new end-to-end test.
- Removed annotate_uncles: a second countable implementation that
production never called, while carrying most of the selection test
coverage. Tests now drive select_uncles_at_production through an
annotate_via_production replay helper -- same assertions, live path.
- Added tests for the two previously uncovered branches of the live
selection: the pmin/below chain walk that resolves parent-on-chain
for candidates whose parent sits below the window, and the
occupied-slot exclusion built from the chain walk.
- theory.q_effective and theory.window_miss_prob were unused and
untested. Now used (the prediction figure reconstructs q_u through
the identity the report quotes) and tested. The window_miss_prob test
records that its "~ e^-W" docstring is the f->0 limit: the true decay
is e^-1.017W at f=1/30, 16% off by W=10.
- Shared sem()/recovery_rate() moved into figures_pernode.py; fig30 and
fig33 regenerated with SEM error bars and the U=0 control curve.
- Fixed the pre-existing E501 in bootstrap_dynamics.py; ruff clean.
Report prose reworked to read standalone: the countable model is
described as the rules under analysis and the former model as a
labelled "unrestricted" comparison baseline, with no dated banners and
no round-to-round narration.
Tests: 209 passed (was 202).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 20:48:38 +02:00
|
|
|
|
print(f"wrote {len(written)} files -> {out}")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
if __name__ == "__main__":
|
|
|
|
|
|
main()
|