The simulator gitignores its runs/ directory, so every table and figure in reports/blend/pd rested on data that existed only on one machine. This adds the sweep outputs of record under reports/blend/pd/data -- one directory per study, 1 MB total -- so any number can be checked against its source, or challenged, without re-running hours of compute. report_numbers.py comes with them: run it and it prints every value the report quotes together with its across-topology standard error, straight from these parquets. It reproduces the report tables exactly. Kept: default (8000 rounds/cell), redundancy (9600), percolation (6400), correlated-churn (6400), fullscale (192, the deliberately lighter 1e6 check). Omitted: the smoke runs, and an earlier 144-rounds/cell redundancy grid whose sampling error produced a non-monotonic delivery curve -- superseded, and the reason the kept grid samples 9600. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Evidence of record
The sweep outputs behind every number in the report. The simulator does not commit
its own runs/ directory — these are the copies of record, kept so that any figure or table can be
re-derived, or challenged, without re-running hours of compute.
Each run directory holds the three tables the simulator writes: propagation.parquet,
adversary.parquet and deanon.parquet.
| directory | config | sampling | backs |
|---|---|---|---|
default/ |
configs/default.yaml |
1 000 rounds × 8 seeds = 8 000/cell | §3.1–§3.5 — delay, observation, eclipse, deanonymization, delivery, coverage |
redundancy/ |
configs/redundancy.yaml |
1 200 × 8 = 9 600/cell | §3.8 — messaging redundancy R = 1…4 |
percolation/ |
configs/percolation.yaml |
800 × 8 = 6 400/cell | §3.5 — the churn threshold u_c = 1 − 1/(degree − 1) |
correlated-churn/ |
configs/correlated-churn.yaml |
800 × 8 = 6 400/cell | §3.9 — correlated AS/region outages vs uniform churn |
fullscale/ |
configs/fullscale.yaml |
64 × 3 = 192/cell | §5 — the 10⁶ scaling check (deliberately lighter; not a source of headline numbers) |
The linkability results (§3.6–§3.7) and both deanonymization rates are closed forms over these
tables rather than separate measurements, so they have no run of their own — pd.linkability
derives them and make verify checks them against Monte-Carlo.
Regenerating the report's numbers
python report_numbers.py
prints every quoted value with its across-topology standard error, straight from the parquets here.
That is the fastest way to check a table in the report against its evidence. It takes optional
paths (report_numbers.py <default> <redundancy> <percolation>) if you want to point it at fresh
runs instead.
Regenerating the data itself
From tools/simulators/blend/pd: make sweep,
make redundancy, make percolation, make correlated-churn, make sweep-fullscale. Results
land in that simulator's runs/<timestamp>_<label>/. Note that the seed streams depend on the
configuration, so re-running reproduces the statistics, not bit-identical numbers, unless the
config is unchanged — in which case it does reproduce exactly.
Two runs from the same session are deliberately not kept: the smoke runs (throwaway, far too
noisy to interpret) and an earlier 144-rounds/cell redundancy grid that was superseded because its
sampling error produced a non-monotonic delivery curve — the reason redundancy/ samples 9 600.