The study started as a peering-degree question and grew well past it: propagation, adversary exposure, deanonymization and time-to-link, reliability under uniform and correlated churn, messaging redundancy, and cover traffic. The pd name no longer describes it. tools/simulators/blend/pd/ -> tools/simulators/blend/, package src/pd -> src/blend, and reports/blend/pd/ -> reports/blend/. Moved with git mv so history follows. The text substitutions are deliberately narrow. pd is also the conventional pandas alias, and pandas genuinely has a pd.plotting submodule, so a blanket pd. -> blend. rewrite would have corrupted four files. Only package-unambiguous forms were changed: from pd.X, -m pd.X, pd.<our module>, PD_BYTES_BUDGET, src/pd, and the pyproject name. All four import pandas as pd lines are untouched and verified. Both READMEs reframed: peering degree is now presented as the primary axis that ties the others together rather than as the subject, and the relative links, which lost a directory level in the move, are corrected. Verified after the move: ruff clean, 101 tests, 45 verify anchors, make targets, the script shims, an end-to-end smoke run, and data/report_numbers.py still reproducing the report tables from the checked-in evidence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.9 KiB
blend — a Monte-Carlo simulator for the Blend network
Measures the Blend network on a seeded peer graph: propagation, adversary exposure, deanonymization, reliability under churn, messaging redundancy and cover traffic. Peering degree is the primary study axis and the one that ties the rest together — it trades off, simultaneously:
- propagation speed — the full delay (ms) of a message: a random sender routes it along a
blend_hops-relay Blend path (each relay a free-running timed-release mix node) and the last relay floods the whole network; - adversary exposure — with a fraction
f_advof adversarial nodes, how many honest nodes are peered with ≥1 adversary (observed) and how many are fully surrounded (eclipsed); - deanonymization — tying propagation to the adversary: how often a message's whole blend path is adversarial (deanonymization — the adversary owns the cascade end-to-end) and how often the honest sender is additionally peered with an adversary (full deanonymization — the message is tied back to its originator); and
- reliability under churn — with a fraction
unresponsive_fracof nodes that relay nothing, the message success-delivery-rate (fraction of messages that survive the whole blend cascade to a responsive final relay) and the flood coverage of those that do.
The peer graph is a seeded random d-regular graph (exactly degree symmetric peers, identical
for everyone from one global seed). This is static-graph analysis — no consensus — so it is far
lighter than the TSI simulators and scales to 10⁶ nodes (sparse CSR + sampled Dijkstra; the
adversary metrics are exact at every N).
Model (all delays in ms)
- Link delay: geographic base (metro 15 → antipodal 200 ms) + exponential transport jitter.
- Processing lag: each node draws a fixed lag from a categorical distribution (default {10, 50, 100} ms at {0.5, 0.4, 0.1}), incurred every time it relays.
- Blend mixing: each relay releases on a free-running clock whose successive intervals are
Uniform{0…
max_blend_delay} whole seconds; a held message waits for the relay's next release (the renewal residual). Mixing happens only at theblend_hopsrelays; the final flood is plain. - Unresponsive nodes: a random
unresponsive_fracof the population relays nothing (its outgoing edges are removed). Relays are drawn from the whole node list blind to responsiveness, so a message dies if any relay on its path is unresponsive — the delivery-rate then tracks(1−unresponsive_frac)^blend_hops. Unresponsive nodes still receive, but they are routing holes, so a delivered flood can strand pockets; a higher peering degree supplies redundant paths that keep coverage high. This axis affects propagation only, not the adversary metrics. - Deanonymization: relays are picked blind to who is adversarial, so P(the whole blend path is
adversarial) is the exact hypergeometric
C(n_adv, blend_hops) / C(N−1, blend_hops)≈f_adv^blend_hops(deanon_rate) — placement-independent, driven by path length, not degree. Multiplying by the fraction of honest nodes with ≥1 adversary peer (observed_frac, which the worst-case-coverage placement maximizes) gives full_deanon_rate — the honest sender is also directly exposed, so the message is tied to its originator. Lengthening the blend path is the dominant defence; a higher degree speeds propagation but raises the chance a sender directly touches the adversary. Both are exact at every N (no Monte-Carlo), like the other adversary metrics. - Messaging redundancy:
redundancyR sends each emission over R independent blend cascades. A node receives the message from whichever cascade reaches it first (arrival times are combined element-wise), so it is delivered if any cascade delivers (delivery = 1−(1−(1−u)^blend_hops)^R) and captured if any cascade is whole-path-adversarial (deanon = 1−(1−f_adv^blend_hops)^R) — the same1−(1−x)^Rlaw, so redundancy trades reliability against anonymity. It buys no extra coverage: a cascade only delivers if the sender could route to its relay, so every delivered cascade floods the sender's own component. R = 1 is the plain single-cascade model (default), to which the whole aggregation reduces exactly. - Churn percolation: the flood only crosses responsive nodes, so it lives on the responsive
sub-graph — site percolation on a d-regular graph, whose giant component survives only while the
responsive fraction exceeds
1/(degree−1). A network tolerates churn up tou_c = 1 − 1/(degree−1)(degree 3 → 0.5, degree 6 → 0.8, degree 16 → 0.93) and shatters above it;configs/percolation.yamlwalks u across the threshold andmake verifychecks it. - Correlated outages:
n_regionssplits the network into equal-sized failure domains andregion_localityplaces that share of each node's peers inside its own domain (the locality matchings keep the graph exactly d-regular).churn_mode: regionalthen fails whole domains instead of scattered nodes, at an identical dead-node count. Locality is what makes this differ from uniform churn at all — with region-blind peering, dropping whole regions removes a uniformly random node set. Clustered failure leaves the survivors fully connected (frac_reached_livestays ~1) while stranding the dead domains (frac_reachedfalls); seeconfigs/correlated-churn.yaml. Regions are failure and peering domains only — link latency does not depend on them. - Linkability over time (
blend.linkability): given the rates above and an emission cadence (one node emits per 30 s slot, chosen ∝ stake), the module derives the time to link an emitter (≈ 30 s·ln(1/(1−α))/(stake·q), inverse in stake) and the time to learn its stake to a threshold from the count of attributable observations. Seeconfigs/redundancy.yamland the report.
Quick start
make install # or reuse a sibling venv: PYTHONPATH=src <python> -m blend.sweep ...
make smoke # fast end-to-end (every code path + 19 of 21 figure builders)
make verify # analytic checks (closed forms + graph invariants)
make test # unit tests
make sweep # configs/default.yaml (N up to 1e5, both adversary modes)
make sweep-fullscale # configs/fullscale.yaml (N up to 1e6, random-mode exact)
make redundancy # configs/redundancy.yaml (R=1..4: delivery vs deanonymization)
make percolation # configs/percolation.yaml (churn threshold u_c = 1-1/(degree-1))
make figures RUN=runs/<dir>
Outputs
Three parquets per run: propagation.parquet (full_delay_ms_*, delivery_rate, coverN_ms, and
two coverage columns — frac_reached over all nodes and frac_reached_live over the responsive
ones, which diverge under correlated churn — vs degree / blend_hops / N / unresponsive_frac /
churn_mode / n_regions / region_locality / redundancy), adversary.parquet
(observed_frac / eclipsed_frac vs degree / f_adv / mode, random + worst-case envelope), and
deanon.parquet (deanon_rate / full_deanon_rate vs degree / blend_hops / redundancy / f_adv /
mode — propagation paths crossed with the adversary set). Figures render all three: delay vs degree /
path length / N, observation and eclipse vs f_adv and degree, delivery and coverage vs the
unresponsive fraction (with the u_c = 1-1/(degree-1) threshold), the deanonymization rates, and the
time-to-link / stake-inference / redundancy curves derived from them by linkability.
Layout
src/blend/: graph (matching-union CSR d-regular), propagation (Blend cascade), mixclock
(release-clock residual), adversary (exact observation/eclipse + deanonymization + placement),
config/engine/sweep/metrics, plotting. configs/ sweeps, tests/, scripts/ shims.
Reports of record live outside the sim at reports/blend/.