Marcin Pawlowski 81c48a38ab
pd review: model completeness, stale numbers, figure coherence
Correctness/completeness pass over the blend material only (TSI untouched).

- report Model section (2) was missing two of the six axes: messaging
  redundancy (R cascades, first-arrival combination) and the emission/linking
  model (30 s stake-proportional cadence, what counts as linked) were defined
  only inline in the findings;
- method note still claimed 200 rounds x 8 topologies, contradicting the 1000
  x 8 the tables now come from;
- design guidance carried two superseded numbers: worst-case observation as
  "+0.15 absolute" (it saturates at 1.000 at degree 8, f_adv 0.2) and the
  redundancy example (0.34 -> 0.72, measured 0.342 -> 0.713);
- figure references were incoherent: Figs 2 and 14 were cited in the text but
  never shown, and Fig 8 was shown but never cited. All 15 embedded figures are
  now cited and all citations resolve;
- simulator README listed two parquets for smoke (there are three) and omitted
  redundancy from the propagation/deanon column lists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 17:59:56 +02:00

90 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# pd — peering-degree Monte-Carlo graph simulator
Quantifies how a node's **peering degree** trades off, in the Blend network:
- **propagation speed** — the full delay (ms) of a message: a random sender routes it along a
`blend_hops`-relay Blend path (each relay a free-running timed-release mix node) and the last
relay floods the whole network;
- **adversary exposure** — with a fraction `f_adv` of adversarial nodes, how many honest nodes are
peered with ≥1 adversary (**observed**) and how many are fully surrounded (**eclipsed**);
- **deanonymization** — tying propagation to the adversary: how often a message's *whole* blend path
is adversarial (**deanonymization** — the adversary owns the cascade end-to-end) and how often the
honest sender is *additionally* peered with an adversary (**full deanonymization** — the message is
tied back to its originator); and
- **reliability under churn** — with a fraction `unresponsive_frac` of nodes that relay nothing, the
**message success-delivery-rate** (fraction of messages that survive the whole blend cascade to a
responsive final relay) and the flood **coverage** of those that do.
The peer graph is a seeded random **d-regular** graph (exactly `degree` symmetric peers, identical
for everyone from one global seed). This is static-graph analysis — no consensus — so it is far
lighter than the TSI simulators and scales to **10⁶ nodes** (sparse CSR + sampled Dijkstra; the
adversary metrics are exact at every N).
## Model (all delays in ms)
- **Link delay:** geographic base (metro 15 → antipodal 200 ms) + exponential transport jitter.
- **Processing lag:** each node draws a fixed lag from a categorical distribution (default
{10, 50, 100} ms at {0.5, 0.4, 0.1}), incurred every time it relays.
- **Blend mixing:** each relay releases on a free-running clock whose successive intervals are
Uniform{0…`max_blend_delay`} whole seconds; a held message waits for the relay's next release
(the renewal residual). Mixing happens only at the `blend_hops` relays; the final flood is plain.
- **Unresponsive nodes:** a random `unresponsive_frac` of the population relays nothing (its outgoing
edges are removed). Relays are drawn from the whole node list *blind to responsiveness*, so a
message dies if any relay on its path is unresponsive — the delivery-rate then tracks
`(1unresponsive_frac)^blend_hops`. Unresponsive nodes still *receive*, but they are routing holes,
so a delivered flood can strand pockets; a higher peering degree supplies redundant paths that keep
coverage high. This axis affects propagation only, not the adversary metrics.
- **Deanonymization:** relays are picked *blind to who is adversarial*, so P(the whole blend path is
adversarial) is the exact hypergeometric `C(n_adv, blend_hops) / C(N1, blend_hops)`
`f_adv^blend_hops` (**deanon_rate**) — placement-independent, driven by path length, not degree.
Multiplying by the fraction of honest nodes with ≥1 adversary peer (`observed_frac`, which the
worst-case-coverage placement maximizes) gives **full_deanon_rate** — the honest sender is *also*
directly exposed, so the message is tied to its originator. Lengthening the blend path is the
dominant defence; a higher degree speeds propagation but *raises* the chance a sender directly
touches the adversary. Both are exact at every N (no Monte-Carlo), like the other adversary metrics.
- **Messaging redundancy:** `redundancy` R sends each emission over R *independent* blend cascades.
A node receives the message from whichever cascade reaches it first (arrival times are combined
element-wise), so it is delivered if any cascade delivers (`delivery = 1(1(1u)^blend_hops)^R`)
and captured if any cascade is whole-path-adversarial (`deanon = 1(1f_adv^blend_hops)^R`) — the
same `1(1x)^R` law, so redundancy trades reliability against anonymity. It buys **no** extra
coverage: a cascade only delivers if the sender could route to its relay, so every delivered
cascade floods the sender's own component. R = 1 is the plain single-cascade model (default), to
which the whole aggregation reduces exactly.
- **Churn percolation:** the flood only crosses responsive nodes, so it lives on the responsive
sub-graph — site percolation on a d-regular graph, whose giant component survives only while the
responsive fraction exceeds `1/(degree1)`. A network tolerates churn up to
`u_c = 1 1/(degree1)` (degree 3 → 0.5, degree 6 → 0.8, degree 16 → 0.93) and shatters above it;
`configs/percolation.yaml` walks u across the threshold and `make verify` checks it.
- **Linkability over time** (`pd.linkability`): given the rates above and an emission cadence (one
node emits per 30 s slot, chosen ∝ stake), the module derives the *time to link* an emitter
(`≈ 30 s·ln(1/(1α))/(stake·q)`, inverse in stake) and the *time to learn its stake* to a threshold
from the count of attributable observations. See `configs/redundancy.yaml` and the report.
## Quick start
```
make install # or reuse a sibling venv: PYTHONPATH=src <python> -m pd.sweep ...
make smoke # fast end-to-end (every code path + 19 of 21 figure builders)
make verify # analytic checks (closed forms + graph invariants)
make test # unit tests
make sweep # configs/default.yaml (N up to 1e5, both adversary modes)
make sweep-fullscale # configs/fullscale.yaml (N up to 1e6, random-mode exact)
make redundancy # configs/redundancy.yaml (R=1..4: delivery vs deanonymization)
make percolation # configs/percolation.yaml (churn threshold u_c = 1-1/(degree-1))
make figures RUN=runs/<dir>
```
## Outputs
Three parquets per run: `propagation.parquet` (`full_delay_ms_*`, `delivery_rate`, `frac_reached`,
`coverN_ms` vs degree / blend_hops / N / unresponsive_frac / redundancy), `adversary.parquet`
(`observed_frac` / `eclipsed_frac` vs degree / f_adv / mode, random + worst-case envelope), and
`deanon.parquet` (`deanon_rate` / `full_deanon_rate` vs degree / blend_hops / redundancy / f_adv /
mode — propagation paths crossed with the adversary set). Figures render all three: delay vs degree /
path length / N, observation and eclipse vs f_adv and degree, delivery and coverage vs the
unresponsive fraction (with the `u_c = 1-1/(degree-1)` threshold), the deanonymization rates, and the
time-to-link / stake-inference / redundancy curves derived from them by `linkability`.
## Layout
`src/pd/`: `graph` (matching-union CSR d-regular), `propagation` (Blend cascade), `mixclock`
(release-clock residual), `adversary` (exact observation/eclipse + deanonymization + placement),
`config`/`engine`/`sweep`/`metrics`, `plotting`. `configs/` sweeps, `tests/`, `scripts/` shims.
Reports of record live outside the sim at `reports/blend/pd/`.