Marcin Pawlowski 36c16a0a98
blend: align section 3.11 with its committed evidence, and expose every study via make
The 3.11 table carried numbers from the ad-hoc analysis that preceded the sweep.
Replaced with the values the checked-in run actually produces (MAP success
0.993/0.905/0.683 clock, 0.989/0.832/0.550 jitter), so every figure in the report
is traceable to data/. The minimum-interval control likewise now quotes the
committed 10.14s vs 10.22s and 0.858 vs 0.860.

cover-traffic was the only study without a make target, and correlated-churn,
cover-traffic and timing were missing from the simulator quick-start. Added.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 18:00:01 +02:00

106 lines
8.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# blend — a Monte-Carlo simulator for the Blend network
Measures the Blend network on a seeded peer graph: propagation, adversary exposure,
deanonymization, reliability under churn, messaging redundancy and cover traffic.
Peering degree is the primary study axis and the one that ties the rest together —
it trades off, simultaneously:
- **propagation speed** — the full delay (ms) of a message: a random sender routes it along a
`blend_hops`-relay Blend path (each relay a free-running timed-release mix node) and the last
relay floods the whole network;
- **adversary exposure** — with a fraction `f_adv` of adversarial nodes, how many honest nodes are
peered with ≥1 adversary (**observed**) and how many are fully surrounded (**eclipsed**);
- **deanonymization** — tying propagation to the adversary: how often a message's *whole* blend path
is adversarial (**deanonymization** — the adversary owns the cascade end-to-end) and how often the
honest sender is *additionally* peered with an adversary (**full deanonymization** — the message is
tied back to its originator); and
- **reliability under churn** — with a fraction `unresponsive_frac` of nodes that relay nothing, the
**message success-delivery-rate** (fraction of messages that survive the whole blend cascade to a
responsive final relay) and the flood **coverage** of those that do.
The peer graph is a seeded random **d-regular** graph (exactly `degree` symmetric peers, identical
for everyone from one global seed). This is static-graph analysis — no consensus — so it is far
lighter than the TSI simulators and scales to **10⁶ nodes** (sparse CSR + sampled Dijkstra; the
adversary metrics are exact at every N).
## Model (all delays in ms)
- **Link delay:** geographic base (metro 15 → antipodal 200 ms) + exponential transport jitter.
- **Processing lag:** each node draws a fixed lag from a categorical distribution (default
{10, 50, 100} ms at {0.5, 0.4, 0.1}), incurred every time it relays.
- **Blend mixing:** each relay releases on a free-running clock whose successive intervals are
Uniform{0…`max_blend_delay`} whole seconds; a held message waits for the relay's next release
(the renewal residual). Mixing happens only at the `blend_hops` relays; the final flood is plain.
- **Unresponsive nodes:** a random `unresponsive_frac` of the population relays nothing (its outgoing
edges are removed). Relays are drawn from the whole node list *blind to responsiveness*, so a
message dies if any relay on its path is unresponsive — the delivery-rate then tracks
`(1unresponsive_frac)^blend_hops`. Unresponsive nodes still *receive*, but they are routing holes,
so a delivered flood can strand pockets; a higher peering degree supplies redundant paths that keep
coverage high. This axis affects propagation only, not the adversary metrics.
- **Deanonymization:** relays are picked *blind to who is adversarial*, so P(the whole blend path is
adversarial) is the exact hypergeometric `C(n_adv, blend_hops) / C(N1, blend_hops)`
`f_adv^blend_hops` (**deanon_rate**) — placement-independent, driven by path length, not degree.
Multiplying by the fraction of honest nodes with ≥1 adversary peer (`observed_frac`, which the
worst-case-coverage placement maximizes) gives **full_deanon_rate** — the honest sender is *also*
directly exposed, so the message is tied to its originator. Lengthening the blend path is the
dominant defence; a higher degree speeds propagation but *raises* the chance a sender directly
touches the adversary. Both are exact at every N (no Monte-Carlo), like the other adversary metrics.
- **Messaging redundancy:** `redundancy` R sends each emission over R *independent* blend cascades.
A node receives the message from whichever cascade reaches it first (arrival times are combined
element-wise), so it is delivered if any cascade delivers (`delivery = 1(1(1u)^blend_hops)^R`)
and captured if any cascade is whole-path-adversarial (`deanon = 1(1f_adv^blend_hops)^R`) — the
same `1(1x)^R` law, so redundancy trades reliability against anonymity. It buys **no** extra
coverage: a cascade only delivers if the sender could route to its relay, so every delivered
cascade floods the sender's own component. R = 1 is the plain single-cascade model (default), to
which the whole aggregation reduces exactly.
- **Churn percolation:** the flood only crosses responsive nodes, so it lives on the responsive
sub-graph — site percolation on a d-regular graph, whose giant component survives only while the
responsive fraction exceeds `1/(degree1)`. A network tolerates churn up to
`u_c = 1 1/(degree1)` (degree 3 → 0.5, degree 6 → 0.8, degree 16 → 0.93) and shatters above it;
`configs/percolation.yaml` walks u across the threshold and `make verify` checks it.
- **Correlated outages:** `n_regions` splits the network into equal-sized **failure domains** and
`region_locality` places that share of each node's peers inside its own domain (the locality
matchings keep the graph exactly d-regular). `churn_mode: regional` then fails whole domains
instead of scattered nodes, at an identical dead-node count. Locality is what makes this differ
from uniform churn at all — with region-blind peering, dropping whole regions removes a uniformly
random node set. Clustered failure leaves the survivors fully connected (`frac_reached_live`
stays ~1) while stranding the dead domains (`frac_reached` falls); see `configs/correlated-churn.yaml`.
Regions are failure and peering domains only — link latency does not depend on them.
- **Linkability over time** (`blend.linkability`): given the rates above and an emission cadence (one
node emits per 30 s slot, chosen ∝ stake), the module derives the *time to link* an emitter
(`≈ 30 s·ln(1/(1α))/(stake·q)`, inverse in stake) and the *time to learn its stake* to a threshold
from the count of attributable observations. See `configs/redundancy.yaml` and the report.
## Quick start
```
make install # or reuse a sibling venv: PYTHONPATH=src <python> -m blend.sweep ...
make smoke # fast end-to-end (every code path + 19 of 21 figure builders)
make verify # analytic checks (closed forms + graph invariants)
make test # unit tests
make sweep # configs/default.yaml (N up to 1e5, both adversary modes)
make sweep-fullscale # configs/fullscale.yaml (N up to 1e6, random-mode exact)
make redundancy # configs/redundancy.yaml (R=1..4: delivery vs deanonymization)
make percolation # configs/percolation.yaml (churn threshold u_c = 1-1/(degree-1))
make correlated-churn # configs/correlated-churn.yaml (AS/region outages vs uniform churn)
make cover-traffic # configs/cover-traffic.yaml (blending, mixing, the stake ceiling)
make timing # configs/timing.yaml (jitter vs clock-tick release under attack)
make figures RUN=runs/<dir>
```
## Outputs
Three parquets per run: `propagation.parquet` (`full_delay_ms_*`, `delivery_rate`, `coverN_ms`, and
two coverage columns — `frac_reached` over all nodes and `frac_reached_live` over the responsive
ones, which diverge under correlated churn — vs degree / blend_hops / N / unresponsive_frac /
`churn_mode` / `n_regions` / `region_locality` / redundancy), `adversary.parquet`
(`observed_frac` / `eclipsed_frac` vs degree / f_adv / mode, random + worst-case envelope), and
`deanon.parquet` (`deanon_rate` / `full_deanon_rate` vs degree / blend_hops / redundancy / f_adv /
mode — propagation paths crossed with the adversary set). Figures render all three: delay vs degree /
path length / N, observation and eclipse vs f_adv and degree, delivery and coverage vs the
unresponsive fraction (with the `u_c = 1-1/(degree-1)` threshold), the deanonymization rates, and the
time-to-link / stake-inference / redundancy curves derived from them by `linkability`.
## Layout
`src/blend/`: `graph` (matching-union CSR d-regular), `propagation` (Blend cascade), `mixclock`
(release-clock residual), `adversary` (exact observation/eclipse + deanonymization + placement),
`config`/`engine`/`sweep`/`metrics`, `plotting`. `configs/` sweeps, `tests/`, `scripts/` shims.
Reports of record live outside the sim at `reports/blend/`.