Marcin Pawlowski e35804f29d
pd: cover traffic -- emission quota and the blending timeline
First half of the cover-traffic work: the two new modules and their tests.

quota.py -- the emission budget. Cover traffic gives every node the same number
of emissions per epoch, which only holds while a node block proposals fit inside
its quota. The bind is exact: alpha_max = ln(1-q)/ln(1-f), where alpha is stake
relative to the INFERRED total D_hat, since that is the denominator the lottery
threshold is derived from. In true stake the ceiling carries the estimator ratio,
s_max = (D_hat/D)*alpha_max, with D_hat/D an input rather than an assumption. The
familiar q/f is a small-q approximation that runs 1.7% high and so overstates the
tolerable stake. Sitting on the mean bind overruns the quota half the time, so
max_alpha_for_confidence gives the ceiling that holds with stated probability.

traffic.py -- the timeline. The rest of the simulator samples independent rounds
and draws each hold from the stationary residual, which has no notion of time and
so can never let two messages meet at a relay. Here every node owns one
free-running clock shared by all messages through it, extended lazily so only the
relays actually visited grow one. A clock sampled once still reproduces
mixclock.mix_wait, so single-message statistics are unchanged.

It separates two quantities that are easy to conflate: mixing (messages a relay
holds at once) and blending (messages it has SEEN between consecutive releases).
Blending is the anonymity set -- every broadcast reaches every node, so an
observer cannot tell which of them the relay forwarded. Gaps sampled at a release
are size-biased, so blending is rate*(2M+1)/3, twice the mean hold, not
rate*M/2 as a naive reading gives. Measured within 1-4% of that at M = 3, 10, 30
and linear in the cover rate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 17:59:57 +02:00
..

pd — peering-degree Monte-Carlo graph simulator

Quantifies how a node's peering degree trades off, in the Blend network:

  • propagation speed — the full delay (ms) of a message: a random sender routes it along a blend_hops-relay Blend path (each relay a free-running timed-release mix node) and the last relay floods the whole network;
  • adversary exposure — with a fraction f_adv of adversarial nodes, how many honest nodes are peered with ≥1 adversary (observed) and how many are fully surrounded (eclipsed);
  • deanonymization — tying propagation to the adversary: how often a message's whole blend path is adversarial (deanonymization — the adversary owns the cascade end-to-end) and how often the honest sender is additionally peered with an adversary (full deanonymization — the message is tied back to its originator); and
  • reliability under churn — with a fraction unresponsive_frac of nodes that relay nothing, the message success-delivery-rate (fraction of messages that survive the whole blend cascade to a responsive final relay) and the flood coverage of those that do.

The peer graph is a seeded random d-regular graph (exactly degree symmetric peers, identical for everyone from one global seed). This is static-graph analysis — no consensus — so it is far lighter than the TSI simulators and scales to 10⁶ nodes (sparse CSR + sampled Dijkstra; the adversary metrics are exact at every N).

Model (all delays in ms)

  • Link delay: geographic base (metro 15 → antipodal 200 ms) + exponential transport jitter.
  • Processing lag: each node draws a fixed lag from a categorical distribution (default {10, 50, 100} ms at {0.5, 0.4, 0.1}), incurred every time it relays.
  • Blend mixing: each relay releases on a free-running clock whose successive intervals are Uniform{0…max_blend_delay} whole seconds; a held message waits for the relay's next release (the renewal residual). Mixing happens only at the blend_hops relays; the final flood is plain.
  • Unresponsive nodes: a random unresponsive_frac of the population relays nothing (its outgoing edges are removed). Relays are drawn from the whole node list blind to responsiveness, so a message dies if any relay on its path is unresponsive — the delivery-rate then tracks (1unresponsive_frac)^blend_hops. Unresponsive nodes still receive, but they are routing holes, so a delivered flood can strand pockets; a higher peering degree supplies redundant paths that keep coverage high. This axis affects propagation only, not the adversary metrics.
  • Deanonymization: relays are picked blind to who is adversarial, so P(the whole blend path is adversarial) is the exact hypergeometric C(n_adv, blend_hops) / C(N1, blend_hops)f_adv^blend_hops (deanon_rate) — placement-independent, driven by path length, not degree. Multiplying by the fraction of honest nodes with ≥1 adversary peer (observed_frac, which the worst-case-coverage placement maximizes) gives full_deanon_rate — the honest sender is also directly exposed, so the message is tied to its originator. Lengthening the blend path is the dominant defence; a higher degree speeds propagation but raises the chance a sender directly touches the adversary. Both are exact at every N (no Monte-Carlo), like the other adversary metrics.
  • Messaging redundancy: redundancy R sends each emission over R independent blend cascades. A node receives the message from whichever cascade reaches it first (arrival times are combined element-wise), so it is delivered if any cascade delivers (delivery = 1(1(1u)^blend_hops)^R) and captured if any cascade is whole-path-adversarial (deanon = 1(1f_adv^blend_hops)^R) — the same 1(1x)^R law, so redundancy trades reliability against anonymity. It buys no extra coverage: a cascade only delivers if the sender could route to its relay, so every delivered cascade floods the sender's own component. R = 1 is the plain single-cascade model (default), to which the whole aggregation reduces exactly.
  • Churn percolation: the flood only crosses responsive nodes, so it lives on the responsive sub-graph — site percolation on a d-regular graph, whose giant component survives only while the responsive fraction exceeds 1/(degree1). A network tolerates churn up to u_c = 1 1/(degree1) (degree 3 → 0.5, degree 6 → 0.8, degree 16 → 0.93) and shatters above it; configs/percolation.yaml walks u across the threshold and make verify checks it.
  • Correlated outages: n_regions splits the network into equal-sized failure domains and region_locality places that share of each node's peers inside its own domain (the locality matchings keep the graph exactly d-regular). churn_mode: regional then fails whole domains instead of scattered nodes, at an identical dead-node count. Locality is what makes this differ from uniform churn at all — with region-blind peering, dropping whole regions removes a uniformly random node set. Clustered failure leaves the survivors fully connected (frac_reached_live stays ~1) while stranding the dead domains (frac_reached falls); see configs/correlated-churn.yaml. Regions are failure and peering domains only — link latency does not depend on them.
  • Linkability over time (pd.linkability): given the rates above and an emission cadence (one node emits per 30 s slot, chosen ∝ stake), the module derives the time to link an emitter (≈ 30 s·ln(1/(1α))/(stake·q), inverse in stake) and the time to learn its stake to a threshold from the count of attributable observations. See configs/redundancy.yaml and the report.

Quick start

make install                 # or reuse a sibling venv: PYTHONPATH=src <python> -m pd.sweep ...
make smoke                   # fast end-to-end (every code path + 19 of 21 figure builders)
make verify                  # analytic checks (closed forms + graph invariants)
make test                    # unit tests
make sweep                   # configs/default.yaml (N up to 1e5, both adversary modes)
make sweep-fullscale         # configs/fullscale.yaml (N up to 1e6, random-mode exact)
make redundancy              # configs/redundancy.yaml (R=1..4: delivery vs deanonymization)
make percolation             # configs/percolation.yaml (churn threshold u_c = 1-1/(degree-1))
make figures RUN=runs/<dir>

Outputs

Three parquets per run: propagation.parquet (full_delay_ms_*, delivery_rate, coverN_ms, and two coverage columns — frac_reached over all nodes and frac_reached_live over the responsive ones, which diverge under correlated churn — vs degree / blend_hops / N / unresponsive_frac / churn_mode / n_regions / region_locality / redundancy), adversary.parquet (observed_frac / eclipsed_frac vs degree / f_adv / mode, random + worst-case envelope), and deanon.parquet (deanon_rate / full_deanon_rate vs degree / blend_hops / redundancy / f_adv / mode — propagation paths crossed with the adversary set). Figures render all three: delay vs degree / path length / N, observation and eclipse vs f_adv and degree, delivery and coverage vs the unresponsive fraction (with the u_c = 1-1/(degree-1) threshold), the deanonymization rates, and the time-to-link / stake-inference / redundancy curves derived from them by linkability.

Layout

src/pd/: graph (matching-union CSR d-regular), propagation (Blend cascade), mixclock (release-clock residual), adversary (exact observation/eclipse + deanonymization + placement), config/engine/sweep/metrics, plotting. configs/ sweeps, tests/, scripts/ shims. Reports of record live outside the sim at reports/blend/pd/.