diff --git a/reports/blend/pd/README.md b/reports/blend/pd/README.md
index 2ff7221..702b5cb 100644
--- a/reports/blend/pd/README.md
+++ b/reports/blend/pd/README.md
@@ -289,9 +289,9 @@ Expressing the cost as *time* rather than as a per-emission probability is what
## 5. Validity and caveats
-- **Structural adversary.** The adversary is modelled as controlling *nodes* and their peerings; it observes messages that traverse relays it owns (deanonymization) and honest nodes it peers (observation). Timing/traffic-analysis correlation across honest relays, and an adversary that adaptively targets the transport path of a *specific* known sender, are outside the model.
+- **Structural adversary; timing correlation is the next study, and is blocked on cover traffic.** The adversary is modelled as controlling *nodes* and their peerings: it observes messages traversing relays it owns (deanonymization) and honest nodes it peers (observation). It does **not** perform timing or traffic-analysis correlation across honest relays. That is not an oversight but a sequencing constraint — a timing adversary is only meaningful against a network that emits **cover traffic**, which the Blend design calls for and this simulator does not yet generate. Modelling cover traffic and then the timing-correlation adversary against it is the natural next step; until then a timing attack here would face an unrealistically bare traffic pattern and its results would not transfer. An adversary that adaptively targets the transport path of a *specific* known sender is likewise outside the current model.
- **Uniform, uncorrelated churn.** Unresponsive nodes are an independent uniform-random sample. Correlated outages (by region/AS) or adversarially placed churn would degrade coverage faster than the uniform percolation of §3.5; the results here are the average, not the worst, case for reliability.
-- **Exactly d-regular topology.** Every node has the same degree. A realistic degree *distribution* (hubs and leaves) would shift both the flood dynamics and the per-node observation exposure; the regular graph is the clean baseline.
+- **Exactly d-regular topology — by design, not by simplification.** Every node has exactly the same number of peers because the protocol requires it: the peer graph is derived by every node from one global seed, so the degree is a protocol constant rather than an emergent property. This is the topology the deployed network will have, so the results are not an idealisation of some heavier-tailed reality — a degree *distribution* would be a different protocol, not a more realistic model of this one.
- **Sampled propagation, exact structure — and what each is worth.** Only the propagation quantities are sampled: they are Monte-Carlo over **1 000 rounds × 8 independent topologies = 8 000 rounds per cell**, which puts the standard error at **≤ 0.009 on every delivery rate**, **≤ 0.001 on every coverage figure** (bar the critical cell below), and **≤ 0.04 s on every full-delay mean** (the redundancy study uses 1 200 × 8 = 9 600 rounds per cell, SEM ≤ 0.006, and the churn-threshold study 800 × 8 = 6 400). That is a digit finer than the tables quote, so the reported two-decimal rates and 0.1-second delays are resolved rather than sampling noise; error bars were computed across topologies, which captures graph-to-graph variation as well as round-to-round. Everything else — the graph invariants, the observation and eclipse counts, both deanonymization rates, and therefore all of §3.6–§3.8's derived times — is closed-form and carries **no sampling error at all** at any N. The worst-case adversary placement is a greedy envelope characterized at N ≤ 10⁵.
- **One cell is intrinsically unstable, by physics rather than sampling.** Coverage at degree 3 with `u = 0.5` sits exactly on that degree's percolation threshold, where the giant component is bimodal: five of eight topologies delivered to no one, three to 0.3–8.6 % of the network. Five times the rounds moved its mean only from 0.019 to 0.024 and left the spread untouched (SEM 0.009), because the variation is across *topologies*, not rounds — it is the critical point. §3.5 therefore states the threshold law rather than a mean there.
- **Measured to 10⁵, not to 10⁶.** Every figure and table here comes from runs at N ≤ 100 000. The simulator is built for 10⁶ — the graph builder, the memory guard and the exact adversary reductions all handle it, and `make sweep-fullscale` runs that grid — but no 10⁶ run backs the numbers in this report. The size-scaling evidence is §3.2's three decades (10³/10⁴/10⁵), over which the full delay rose 18 %; extrapolating that trend to 10⁶ is an inference, not a measurement.
diff --git a/tools/simulators/blend/pd/Makefile b/tools/simulators/blend/pd/Makefile
index 1cb2f5a..343832a 100644
--- a/tools/simulators/blend/pd/Makefile
+++ b/tools/simulators/blend/pd/Makefile
@@ -8,7 +8,7 @@ export OPENBLAS_NUM_THREADS := 1
export MKL_NUM_THREADS := 1
export NUMEXPR_NUM_THREADS := 1
-.PHONY: install smoke sweep sweep-fullscale redundancy percolation figures verify test lint clean
+.PHONY: install smoke sweep sweep-fullscale redundancy percolation correlated-churn figures verify test lint clean
# The stamp is the real install; targets below depend on it so `make sweep` (etc.) auto-installs
# on a fresh checkout and re-installs whenever pyproject.toml changes.
@@ -35,6 +35,9 @@ redundancy: $(STAMP) ## messaging redundancy R=1..4 (delivery vs deanonymization
percolation: $(STAMP) ## churn threshold: coverage collapse at u_c = 1 - 1/(degree-1)
$(PY) -m pd.sweep --config configs/percolation.yaml
+correlated-churn: $(STAMP) ## correlated AS/region outages vs uniform churn, matched fractions
+ $(PY) -m pd.sweep --config configs/correlated-churn.yaml
+
figures: $(STAMP) ## make figures RUN=runs/
$(PY) -m pd.plotting.make_figures --run $(RUN)
diff --git a/tools/simulators/blend/pd/configs/correlated-churn.yaml b/tools/simulators/blend/pd/configs/correlated-churn.yaml
new file mode 100644
index 0000000..f969e9c
--- /dev/null
+++ b/tools/simulators/blend/pd/configs/correlated-churn.yaml
@@ -0,0 +1,29 @@
+# Correlated (AS/region) churn vs uniform churn, at matched churn fractions.
+#
+# Real outages are not independent: a datacentre, AS or region goes dark as a unit. This config
+# partitions the network into `n_regions` failure domains and compares two ways of removing the
+# SAME number of nodes -- `uniform` (scattered, the §3.5 model) against `regional` (whole domains).
+#
+# Correlation only has a structural effect if the peer graph itself is region-aware: with
+# region-blind peering, taking out whole regions removes a uniformly random set of nodes and is
+# therefore indistinguishable from uniform churn. So `region_locality` places 75% of every node's
+# peers inside its own region, which is what makes a failure domain a *connectivity* domain too.
+# Both churn modes run on the SAME topology, so the comparison is controlled.
+#
+# 800 rounds x 8 seeds = 6400 rounds per cell, matching the churn study in configs/percolation.yaml.
+n_nodes: [20000]
+degree: [4, 8, 16]
+blend_hops: [1, 3]
+max_blend_delay: [0]
+unresponsive_frac: [0.0, 0.2, 0.4, 0.5, 0.6, 0.7, 0.8]
+churn_mode: [uniform, regional]
+redundancy: [1]
+f_adv: [0.2]
+adversary_mode: [random]
+seeds: 8
+base:
+ n_regions: 40 # 500 nodes per failure domain
+ region_locality: 0.75 # 3 of every 4 peers inside the node's own region
+ n_rounds: 800
+ n_placements: 1
+ worstcase_max_n: 100000
diff --git a/tools/simulators/blend/pd/src/pd/config.py b/tools/simulators/blend/pd/src/pd/config.py
index dca9c61..7c663b8 100644
--- a/tools/simulators/blend/pd/src/pd/config.py
+++ b/tools/simulators/blend/pd/src/pd/config.py
@@ -11,9 +11,11 @@ LatencyDist = Literal["geo", "fixed", "uniform", "exp"]
AdversaryMode = Literal[
"random", "worstcase_coverage", "worstcase_eclipse", "worstcase_degree"
]
+ChurnMode = Literal["uniform", "regional"]
_DISTS = ("geo", "fixed", "uniform", "exp")
_MODES = ("random", "worstcase_coverage", "worstcase_eclipse", "worstcase_degree")
WORSTCASE_MODES = ("worstcase_coverage", "worstcase_eclipse", "worstcase_degree")
+_CHURN_MODES = ("uniform", "regional")
@dataclass(frozen=True)
@@ -23,11 +25,14 @@ class SimConfig:
# --- network ---
n_nodes: int = 1000 # must be even (matching-union construction)
degree: int = 8 # peering degree — the primary study axis
+ n_regions: int = 1 # failure domains (AS/region); 1 = no regional structure
+ region_locality: float = 0.0 # fraction of a node's peers drawn inside its own region
# --- propagation (Blend cascade, delays in ms) ---
blend_hops: int = 3 # relay-path length (swept)
max_blend_delay: int = 3 # free-running release-clock max interval, whole SECONDS
unresponsive_frac: float = 0.0 # ratio of nodes that do NOT relay any messages (swept)
+ churn_mode: ChurnMode = "uniform" # "uniform" = independent nodes; "regional" = whole regions
redundancy: int = 1 # copies per emission via R independent cascades (swept)
n_rounds: int = 200 # random-sender rounds per topology
transport_jitter_mean_ms: float = 5.0
@@ -61,6 +66,25 @@ class SimConfig:
raise ValueError(f"need 0 <= unresponsive_frac < 1, got {self.unresponsive_frac}")
if self.redundancy < 1:
raise ValueError(f"redundancy must be >= 1, got {self.redundancy}")
+ if self.n_regions < 1:
+ raise ValueError(f"n_regions must be >= 1, got {self.n_regions}")
+ if self.n_regions > 1:
+ # regions must be equal-sized and each internally matchable (even size)
+ if self.n_nodes % self.n_regions != 0:
+ raise ValueError(
+ f"n_nodes ({self.n_nodes}) must divide evenly into "
+ f"n_regions ({self.n_regions})")
+ if (self.n_nodes // self.n_regions) % 2 != 0:
+ raise ValueError(
+ f"region size ({self.n_nodes // self.n_regions}) must be even")
+ if not (0.0 <= self.region_locality <= 1.0):
+ raise ValueError(f"need 0 <= region_locality <= 1, got {self.region_locality}")
+ if self.region_locality > 0.0 and self.n_regions < 2:
+ raise ValueError("region_locality > 0 requires n_regions >= 2")
+ if self.churn_mode not in _CHURN_MODES:
+ raise ValueError(f"churn_mode must be one of {_CHURN_MODES}")
+ if self.churn_mode == "regional" and self.n_regions < 2:
+ raise ValueError("churn_mode 'regional' requires n_regions >= 2")
if not (0.0 <= self.f_adv < 1.0):
raise ValueError(f"need 0 <= f_adv < 1, got {self.f_adv}")
if self.n_rounds < 1 or self.n_placements < 1:
@@ -88,8 +112,9 @@ class SimConfig:
def key(self) -> tuple:
"""Hashable identity used to seed RNGs deterministically."""
return (
- self.n_nodes, self.degree, self.blend_hops, self.max_blend_delay,
- self.unresponsive_frac, self.redundancy, self.n_rounds,
+ self.n_nodes, self.degree, self.n_regions, self.region_locality,
+ self.blend_hops, self.max_blend_delay,
+ self.unresponsive_frac, self.churn_mode, self.redundancy, self.n_rounds,
self.transport_jitter_mean_ms, self.processing_lags_ms, self.processing_lag_probs,
self.link_latency_dist, self.link_latency_mean_ms, self.coverage_pcts,
self.f_adv, self.adversary_mode, self.n_placements, self.worstcase_max_n,
@@ -110,6 +135,7 @@ class SweepConfig:
blend_hops: list[int] = field(default_factory=lambda: [3])
max_blend_delay: list[int] = field(default_factory=lambda: [3])
unresponsive_frac: list[float] = field(default_factory=lambda: [0.0])
+ churn_mode: list[str] = field(default_factory=lambda: ["uniform"])
redundancy: list[int] = field(default_factory=lambda: [1])
f_adv: list[float] = field(default_factory=lambda: [0.1, 0.2, 0.33, 0.5])
adversary_mode: list[str] = field(default_factory=lambda: ["random"])
@@ -148,7 +174,7 @@ class SweepConfig:
d = dict(d)
base = d.pop("base", {})
known = {"n_nodes", "degree", "blend_hops", "max_blend_delay", "unresponsive_frac",
- "redundancy", "f_adv", "adversary_mode", "seeds"}
+ "churn_mode", "redundancy", "f_adv", "adversary_mode", "seeds"}
unknown = set(d) - known
if unknown:
raise ValueError(f"unknown sweep keys: {sorted(unknown)}")
diff --git a/tools/simulators/blend/pd/src/pd/engine.py b/tools/simulators/blend/pd/src/pd/engine.py
index e7e85d7..0d24443 100644
--- a/tools/simulators/blend/pd/src/pd/engine.py
+++ b/tools/simulators/blend/pd/src/pd/engine.py
@@ -20,6 +20,7 @@ from .rng import placement_seedseq, responsive_seedseq, round_seedseq
def run_graph_cell(base: SimConfig, prop_grid: list[tuple[int, int]],
unresponsive_fracs: list[float], redundancies: list[int],
adv_grid: list[tuple[float, str]],
+ churn_modes: list[str] | None = None,
) -> tuple[list[dict], list[dict], list[dict]]:
"""Build ``base``'s topology once; return (propagation, adversary, deanonymization rows).
@@ -32,18 +33,22 @@ def run_graph_cell(base: SimConfig, prop_grid: list[tuple[int, int]],
"""
graph = build_graph(base)
blend_hops_set = sorted({bh for bh, _ in prop_grid})
+ modes = churn_modes or [base.churn_mode]
prop_rows: list[dict] = []
- for uf in unresponsive_fracs:
- responsive = assign_responsive(
- base.n_nodes, uf, np.random.default_rng(responsive_seedseq(base, uf)))
- for blend_hops, max_blend_delay in prop_grid:
- for R in redundancies:
- rng = np.random.default_rng(
- round_seedseq(base, blend_hops, max_blend_delay, uf, R))
- prop = propagation_metrics(
- graph, blend_hops, max_blend_delay, uf, R, responsive, base, rng)
- prop_rows.append(propagation_row(base, blend_hops, max_blend_delay, uf, R, prop))
+ for cm in modes:
+ for uf in unresponsive_fracs:
+ responsive = assign_responsive(
+ base.n_nodes, uf, np.random.default_rng(responsive_seedseq(base, uf, cm)),
+ cm, base.n_regions)
+ for blend_hops, max_blend_delay in prop_grid:
+ for R in redundancies:
+ rng = np.random.default_rng(
+ round_seedseq(base, blend_hops, max_blend_delay, uf, R))
+ prop = propagation_metrics(
+ graph, blend_hops, max_blend_delay, uf, R, responsive, base, rng)
+ prop_rows.append(
+ propagation_row(base, blend_hops, max_blend_delay, uf, R, prop, cm))
adv_rows: list[dict] = []
deanon_rows: list[dict] = []
@@ -70,7 +75,9 @@ def run_trajectory(config: SimConfig) -> dict:
graph = build_graph(config)
uf = config.unresponsive_frac
responsive = assign_responsive(
- config.n_nodes, uf, np.random.default_rng(responsive_seedseq(config, uf)))
+ config.n_nodes, uf,
+ np.random.default_rng(responsive_seedseq(config, uf, config.churn_mode)),
+ config.churn_mode, config.n_regions)
R = config.redundancy
prng = np.random.default_rng(
round_seedseq(config, config.blend_hops, config.max_blend_delay, uf, R))
diff --git a/tools/simulators/blend/pd/src/pd/graph.py b/tools/simulators/blend/pd/src/pd/graph.py
index 9f0b2bb..042d924 100644
--- a/tools/simulators/blend/pd/src/pd/graph.py
+++ b/tools/simulators/blend/pd/src/pd/graph.py
@@ -45,10 +45,23 @@ class Graph:
return csr_matrix((data, self.indices, self.indptr), shape=(self.n, self.n))
-def build_regular_edges(n: int, degree: int, rng: np.random.Generator) -> np.ndarray:
+def region_of(n: int, n_regions: int) -> np.ndarray:
+ """Region (failure-domain) id per node: equal-sized contiguous blocks."""
+ if n_regions <= 1:
+ return np.zeros(n, dtype=np.int64)
+ return np.arange(n, dtype=np.int64) // (n // n_regions)
+
+
+def build_regular_edges(n: int, degree: int, rng: np.random.Generator,
+ n_regions: int = 1, region_locality: float = 0.0) -> np.ndarray:
"""Exactly d-regular simple undirected edge list, shape (E, 2) with u < v.
Requires ``n`` even and ``1 <= degree < n``. Deterministic in ``rng``.
+
+ With ``region_locality > 0``, ``round(locality * degree)`` of the ``degree`` matchings are drawn
+ *within* each region instead of globally, so a node keeps that share of its peers inside its own
+ failure domain. Each region is matched independently (its size must be even), so the union is
+ still exactly d-regular -- locality changes *where* the peers are, never how many.
"""
if n % 2 != 0:
raise ValueError("n must be even")
@@ -57,9 +70,20 @@ def build_regular_edges(n: int, degree: int, rng: np.random.Generator) -> np.nda
h = n // 2
lo = np.empty(degree * h, dtype=np.int64)
hi = np.empty(degree * h, dtype=np.int64)
+ n_local = int(round(region_locality * degree)) if n_regions > 1 else 0
+ members = ([np.where(region_of(n, n_regions) == r)[0] for r in range(n_regions)]
+ if n_local else [])
for m in range(degree):
- perm = rng.permutation(n)
- a, b = perm[0::2], perm[1::2]
+ if m < n_local: # intra-region matching, region by region
+ parts_a, parts_b = [], []
+ for mem in members:
+ perm = mem[rng.permutation(mem.shape[0])]
+ parts_a.append(perm[0::2])
+ parts_b.append(perm[1::2])
+ a, b = np.concatenate(parts_a), np.concatenate(parts_b)
+ else: # global matching
+ perm = rng.permutation(n)
+ a, b = perm[0::2], perm[1::2]
lo[m * h:(m + 1) * h] = np.minimum(a, b)
hi[m * h:(m + 1) * h] = np.maximum(a, b)
# drop parallel edges (keep first occurrence of each undirected pair)
@@ -165,7 +189,7 @@ def build_graph(config: SimConfig) -> Graph:
check_alloc(int(n * degree * 24), "d-regular CSR (indices+base+src)",
f"N={n}, degree={degree}")
rng = np.random.default_rng(graph_seedseq(config))
- edges = build_regular_edges(n, degree, rng)
+ edges = build_regular_edges(n, degree, rng, config.n_regions, config.region_locality)
base_u = latency.sample_link_latencies(edges.shape[0], config, rng)
p = latency.assign_processing_lags(n, config, rng)
rows = np.concatenate([edges[:, 0], edges[:, 1]])
diff --git a/tools/simulators/blend/pd/src/pd/metrics.py b/tools/simulators/blend/pd/src/pd/metrics.py
index 42a57ea..f049a5e 100644
--- a/tools/simulators/blend/pd/src/pd/metrics.py
+++ b/tools/simulators/blend/pd/src/pd/metrics.py
@@ -6,13 +6,17 @@ from .config import SimConfig
def propagation_row(config: SimConfig, blend_hops: int, max_blend_delay: int,
- unresponsive_frac: float, redundancy: int, prop: dict) -> dict:
+ unresponsive_frac: float, redundancy: int, prop: dict,
+ churn_mode: str | None = None) -> dict:
return {
"n_nodes": config.n_nodes,
"degree": config.degree,
"blend_hops": blend_hops,
"max_blend_delay": max_blend_delay,
"unresponsive_frac": unresponsive_frac,
+ "churn_mode": churn_mode or config.churn_mode,
+ "n_regions": config.n_regions,
+ "region_locality": config.region_locality,
"redundancy": redundancy,
"graph_seed": config.graph_seed,
"n_rounds": config.n_rounds,
diff --git a/tools/simulators/blend/pd/src/pd/plotting/figures.py b/tools/simulators/blend/pd/src/pd/plotting/figures.py
index f2ad05b..06f4cd0 100644
--- a/tools/simulators/blend/pd/src/pd/plotting/figures.py
+++ b/tools/simulators/blend/pd/src/pd/plotting/figures.py
@@ -147,6 +147,40 @@ def coverage_percolation(prop: pd.DataFrame, adv: pd.DataFrame):
return fig
+def churn_correlated_vs_uniform(prop: pd.DataFrame, adv: pd.DataFrame):
+ """Correlated (AS/region) outages against uniform churn at matched churn fractions.
+
+ Solid = coverage of the *live* network (what the survivors still reach), dashed = coverage of
+ all nodes (which counts the dead, who can still receive). Clustered failure leaves survivors
+ better connected but strands whole dead regions, so the two curves separate in opposite
+ directions -- the single number "coverage" hides which of the two you mean.
+ """
+ if (prop is None or not len(prop) or "churn_mode" not in prop
+ or prop["churn_mode"].nunique() < 2):
+ return None
+ import matplotlib.pyplot as plt
+ style.apply_style()
+ n = _largest_n(prop)
+ bh = int(sorted(prop["blend_hops"].unique())[0])
+ d = prop[(prop.n_nodes == n) & (prop.blend_hops == bh)]
+ deg = _median_degree(d)
+ d = d[d.degree == deg]
+ fig, ax = plt.subplots()
+ for i, mode in enumerate(sorted(d.churn_mode.unique())):
+ c = style.color_for(i)
+ s = d[d.churn_mode == mode].groupby("unresponsive_frac").agg(
+ live=("frac_reached_live", "mean"), all_=("frac_reached", "mean")).reset_index()
+ ax.plot(s.unresponsive_frac, s.live, "-o", ms=4, color=c, label=f"{mode}: live network")
+ ax.plot(s.unresponsive_frac, s.all_, "--s", ms=3, color=c, alpha=0.65,
+ label=f"{mode}: all nodes")
+ ax.set_xlabel("unresponsive fraction u")
+ ax.set_ylabel("flood coverage")
+ ax.set_ylim(-0.02, 1.02)
+ ax.set_title(f"Correlated vs uniform churn (N={n:,}, degree={deg}, {bh} hop)")
+ ax.legend(fontsize=7)
+ return fig
+
+
def delay_vs_blendhops(prop: pd.DataFrame, adv: pd.DataFrame):
if prop is None or not len(prop) or prop["blend_hops"].nunique() < 2:
return None
diff --git a/tools/simulators/blend/pd/src/pd/plotting/make_figures.py b/tools/simulators/blend/pd/src/pd/plotting/make_figures.py
index 0e1b42f..e73761a 100644
--- a/tools/simulators/blend/pd/src/pd/plotting/make_figures.py
+++ b/tools/simulators/blend/pd/src/pd/plotting/make_figures.py
@@ -23,6 +23,7 @@ _BUILDERS = [
("10_delivery_vs_unresponsive", figures.delivery_vs_unresponsive),
("11_coverage_vs_unresponsive", figures.coverage_vs_unresponsive),
("20_coverage_percolation", figures.coverage_percolation),
+ ("22_churn_correlated_vs_uniform", figures.churn_correlated_vs_uniform),
]
# (prop, adv, deanon) builders — deanonymization crosses propagation paths with the adversary set,
diff --git a/tools/simulators/blend/pd/src/pd/propagation.py b/tools/simulators/blend/pd/src/pd/propagation.py
index 9bf2ebf..65a9d1a 100644
--- a/tools/simulators/blend/pd/src/pd/propagation.py
+++ b/tools/simulators/blend/pd/src/pd/propagation.py
@@ -33,12 +33,38 @@ from .memguard import check_alloc
from .mixclock import mix_wait
-def assign_responsive(n: int, unresponsive_frac: float, rng: np.random.Generator) -> np.ndarray:
- """Boolean mask (True == responsive/relaying). A random ``unresponsive_frac`` are set False."""
+def assign_responsive(n: int, unresponsive_frac: float, rng: np.random.Generator,
+ churn_mode: str = "uniform", n_regions: int = 1) -> np.ndarray:
+ """Boolean mask (True == responsive/relaying); exactly ``round(frac*n)`` nodes are set False.
+
+ ``uniform`` drops nodes independently at random -- the model of uncorrelated failure.
+ ``regional`` drops whole **failure domains** (AS/region) at a time, the model of a correlated
+ outage: regions are taken down in random order until the quota is met, with the final partial
+ region trimmed at random so the *count* of dead nodes matches ``uniform`` exactly. The two modes
+ are therefore compared at identical churn, differing only in how the failures are arranged.
+ """
responsive = np.ones(n, dtype=bool)
n_unresp = int(round(unresponsive_frac * n))
- if n_unresp > 0:
+ if n_unresp <= 0:
+ return responsive
+ if churn_mode == "uniform" or n_regions <= 1:
responsive[rng.choice(n, size=n_unresp, replace=False)] = False
+ return responsive
+ if churn_mode != "regional":
+ raise ValueError(f"unknown churn_mode {churn_mode!r}")
+ from .graph import region_of
+ region = region_of(n, n_regions)
+ dead = 0
+ for r in rng.permutation(n_regions): # fail whole regions in random order
+ members = np.where(region == r)[0]
+ if dead + members.shape[0] <= n_unresp:
+ responsive[members] = False
+ dead += members.shape[0]
+ else: # partial region: trim to hit the exact count
+ take = n_unresp - dead
+ if take > 0:
+ responsive[rng.choice(members, size=take, replace=False)] = False
+ break
return responsive
@@ -130,6 +156,7 @@ def propagation_metrics(graph: Graph, blend_hops: int, max_blend_delay: int,
("full_delay_ms_mean", "full_delay_ms_p50", "full_delay_ms_p90",
"full_delay_ms_p99", "path_delay_ms_mean", "broadcast_delay_ms_mean")}
out["frac_reached"] = 0.0
+ out["frac_reached_live"] = 0.0
out["delivery_rate"] = 0.0
for pc in pcts:
out[f"cover{int(pc)}_ms"] = float("nan")
@@ -139,7 +166,7 @@ def propagation_metrics(graph: Graph, blend_hops: int, max_blend_delay: int,
return _empty() # no responsive sender, or too few nodes to draw a distinct path
R = int(redundancy)
- fulls, paths, bcasts, fracs = [], [], [], []
+ fulls, paths, bcasts, fracs, fracs_live = [], [], [], [], []
covers = [[] for _ in pcts]
delivered = 0
for _ in range(config.n_rounds):
@@ -166,6 +193,7 @@ def propagation_metrics(graph: Graph, blend_hops: int, max_blend_delay: int,
paths.append(path_min)
bcasts.append(full - path_min)
fracs.append(float(finite.mean()))
+ fracs_live.append(float(finite[responsive].mean())) # coverage of the live network
rel = reached - path_min # coverage measured from the quickest path's release
for j, pc in enumerate(pcts):
covers[j].append(float(np.percentile(rel, pc)))
@@ -185,6 +213,7 @@ def propagation_metrics(graph: Graph, blend_hops: int, max_blend_delay: int,
"path_delay_ms_mean": float(np.mean(paths)),
"broadcast_delay_ms_mean": float(np.mean(bcasts)),
"frac_reached": float(np.mean(fracs)),
+ "frac_reached_live": float(np.mean(fracs_live)),
"delivery_rate": delivery_rate,
}
for pc, col in zip(pcts, covers, strict=True):
diff --git a/tools/simulators/blend/pd/src/pd/rng.py b/tools/simulators/blend/pd/src/pd/rng.py
index 9689c6e..d9dd242 100644
--- a/tools/simulators/blend/pd/src/pd/rng.py
+++ b/tools/simulators/blend/pd/src/pd/rng.py
@@ -36,16 +36,18 @@ def graph_seedseq(config: SimConfig) -> np.random.SeedSequence:
"""Topology-only seed: peer graph + processing lags depend on these fields alone."""
return np.random.SeedSequence(_digest(
config.root_seed, "graph", config.n_nodes, config.degree, config.graph_seed,
+ config.n_regions, config.region_locality,
config.link_latency_dist, config.link_latency_mean_ms,
config.processing_lags_ms, config.processing_lag_probs,
))
-def responsive_seedseq(config: SimConfig, unresponsive_frac: float) -> np.random.SeedSequence:
- """Which nodes are responsive: fixed per (topology, unresponsive_frac), not per round."""
+def responsive_seedseq(config: SimConfig, unresponsive_frac: float,
+ churn_mode: str = "uniform") -> np.random.SeedSequence:
+ """Which nodes are responsive: fixed per (topology, unresponsive_frac, churn_mode)."""
return np.random.SeedSequence(_digest(
config.root_seed, "responsive", config.n_nodes, config.degree, config.graph_seed,
- unresponsive_frac,
+ unresponsive_frac, churn_mode, config.n_regions,
))
diff --git a/tools/simulators/blend/pd/src/pd/sweep.py b/tools/simulators/blend/pd/src/pd/sweep.py
index 071e382..795cb8a 100644
--- a/tools/simulators/blend/pd/src/pd/sweep.py
+++ b/tools/simulators/blend/pd/src/pd/sweep.py
@@ -37,8 +37,10 @@ def new_run_dir(outdir: Path, label: str) -> Path:
return run_dir
-def _cell_worker(base: SimConfig, prop_grid, unresponsive_fracs, redundancies, adv_grid):
- return run_graph_cell(base, prop_grid, unresponsive_fracs, redundancies, adv_grid)
+def _cell_worker(base: SimConfig, prop_grid, unresponsive_fracs, redundancies, adv_grid,
+ churn_modes):
+ return run_graph_cell(base, prop_grid, unresponsive_fracs, redundancies, adv_grid,
+ churn_modes)
def run_sweep(sweep: SweepConfig,
@@ -47,10 +49,12 @@ def run_sweep(sweep: SweepConfig,
prop_grid = sweep.prop_grid()
unresponsive_fracs = list(sweep.unresponsive_frac)
redundancies = list(sweep.redundancy)
+ churn_modes = list(sweep.churn_mode)
adv_grid = sweep.adv_grid()
bases = [sweep.base_config(n, d, g) for (n, d, g) in cells]
results = Parallel(n_jobs=n_jobs, prefer="processes")(
- delayed(_cell_worker)(base, prop_grid, unresponsive_fracs, redundancies, adv_grid)
+ delayed(_cell_worker)(base, prop_grid, unresponsive_fracs, redundancies, adv_grid,
+ churn_modes)
for base in tqdm(bases, desc="topologies")
)
prop_rows = [r for pr, _, _ in results for r in pr]
diff --git a/tools/simulators/blend/pd/src/pd/verify.py b/tools/simulators/blend/pd/src/pd/verify.py
index bd5e97e..39afa89 100644
--- a/tools/simulators/blend/pd/src/pd/verify.py
+++ b/tools/simulators/blend/pd/src/pd/verify.py
@@ -188,6 +188,35 @@ def main(argv: list[str] | None = None) -> int:
ok &= _check(f"percolation d={degree} (u_c={u_c:.2f})", below > 0.1 and above < 0.02,
f"giant {below:.3f} at u_c-0.15 -> {above:.4f} at u_c+0.15")
+ # 8. correlated (regional) churn vs uniform, at an identical number of dead nodes. With most
+ # peers inside the failure domain, losing whole domains leaves the survivors fully connected
+ # -- so every live relay is still routable and delivery equals the live-relay rate -- while
+ # the same number of scattered failures breaks routes and loses delivery below it.
+ n_c, nr_c, u_c2 = 4000, 20, 0.5
+ cc = SimConfig(n_nodes=n_c, degree=4, n_regions=nr_c, region_locality=0.75, blend_hops=1,
+ max_blend_delay=0, transport_jitter_mean_ms=0.0, unresponsive_frac=u_c2,
+ n_rounds=1500, graph_seed=0)
+ gc = build_graph(cc)
+ res = {}
+ for cm in ("uniform", "regional"):
+ mask = assign_responsive(
+ n_c, u_c2, np.random.default_rng(responsive_seedseq(cc, u_c2, cm)), cm, nr_c)
+ ok &= _check(f"churn mode {cm} kills the same count",
+ int((~mask).sum()) == int(round(u_c2 * n_c)),
+ f"{int((~mask).sum())} dead of {n_c}")
+ prng = np.random.default_rng(round_seedseq(cc, 1, 0, u_c2, 1))
+ m = propagation_metrics(gc, 1, 0, u_c2, 1, mask, cc, prng)
+ res[cm] = m
+ ok &= _check("regional churn spares the live network",
+ res["regional"]["frac_reached_live"] > 0.98
+ and res["regional"]["frac_reached_live"] > res["uniform"]["frac_reached_live"],
+ f"live coverage regional {res['regional']['frac_reached_live']:.3f}"
+ f" vs uniform {res['uniform']['frac_reached_live']:.3f}")
+ ok &= _check("uniform churn loses more delivery to broken routes",
+ res["regional"]["delivery_rate"] > res["uniform"]["delivery_rate"],
+ f"delivery regional {res['regional']['delivery_rate']:.3f}"
+ f" vs uniform {res['uniform']['delivery_rate']:.3f}")
+
print("OK" if ok else "FAILURES PRESENT")
return 0 if ok else 1
diff --git a/tools/simulators/blend/pd/tests/test_config.py b/tools/simulators/blend/pd/tests/test_config.py
index 20a756b..c5e5107 100644
--- a/tools/simulators/blend/pd/tests/test_config.py
+++ b/tools/simulators/blend/pd/tests/test_config.py
@@ -8,12 +8,15 @@ from pd.config import SimConfig, SweepConfig
def test_key_covers_every_field():
fields = [f.name for f in dataclasses.fields(SimConfig)]
assert len(SimConfig().key()) == len(fields)
- base = SimConfig()
+ # n_regions=2 in the base so the region/churn fields can each be varied on their own
+ # (region_locality and churn_mode="regional" both require n_regions >= 2)
+ base = SimConfig(n_regions=2)
for name in fields:
cur = getattr(base, name)
- alt = {"n_nodes": 2000, "degree": 4, "blend_hops": 2, "max_blend_delay": 5,
- "unresponsive_frac": 0.2, "redundancy": 2, "n_rounds": 10,
- "transport_jitter_mean_ms": 1.0,
+ alt = {"n_nodes": 2000, "degree": 4, "n_regions": 4, "region_locality": 0.5,
+ "blend_hops": 2, "max_blend_delay": 5,
+ "unresponsive_frac": 0.2, "churn_mode": "regional", "redundancy": 2,
+ "n_rounds": 10, "transport_jitter_mean_ms": 1.0,
"processing_lags_ms": (11.0, 51.0, 101.0), "processing_lag_probs": (0.6, 0.3, 0.1),
"link_latency_dist": "fixed", "link_latency_mean_ms": 1.0,
"coverage_pcts": (25.0,), "f_adv": 0.1, "adversary_mode": "worstcase_coverage",
diff --git a/tools/simulators/blend/pd/tests/test_propagation.py b/tools/simulators/blend/pd/tests/test_propagation.py
index cc3907a..1fc6546 100644
--- a/tools/simulators/blend/pd/tests/test_propagation.py
+++ b/tools/simulators/blend/pd/tests/test_propagation.py
@@ -175,7 +175,9 @@ def test_redundancy_buys_no_coverage_even_when_fragmented():
def test_redundant_cascades_flood_the_same_component():
"""Direct check of the mechanism: with several cascades delivered in one round, the union of
their reached sets equals the largest single one."""
- n, u = 4000, 0.5
+ # u just below degree 3's percolation threshold (0.5): the graph is thinned and lossy, but
+ # deliveries are still common enough that the multi-cascade case actually arises.
+ n, u = 4000, 0.4
cfg = SimConfig(n_nodes=n, degree=3, blend_hops=1, max_blend_delay=0,
transport_jitter_mean_ms=0.0, unresponsive_frac=u, graph_seed=0)
g = build_graph(cfg)
@@ -198,3 +200,59 @@ def test_redundant_cascades_flood_the_same_component():
union = np.logical_or.reduce(masks)
assert int(union.sum()) == max(int(m.sum()) for m in masks)
assert checked > 0 # the multi-delivery case did occur
+
+
+# --- regional (correlated) churn -----------------------------------------------------------------
+
+def test_regional_churn_drops_whole_regions_and_matches_the_uniform_count():
+ """Correlated churn kills failure domains, not scattered nodes -- at the same total count."""
+ from pd.graph import region_of
+ n, n_regions, u = 1000, 10, 0.3
+ rng = np.random.default_rng(0)
+ mask = assign_responsive(n, u, rng, "regional", n_regions)
+ assert int((~mask).sum()) == 300 # exactly the same quota as uniform
+ region = region_of(n, n_regions)
+ dead_per_region = [int((~mask[region == r]).sum()) for r in range(n_regions)]
+ # every region is either wholly dead (100) or wholly alive (0), bar at most one trimmed region
+ partial = [d for d in dead_per_region if 0 < d < 100]
+ assert len(partial) <= 1
+ assert sum(1 for d in dead_per_region if d == 100) == 3
+
+
+def test_uniform_churn_scatters_across_all_regions():
+ from pd.graph import region_of
+ n, n_regions, u = 1000, 10, 0.3
+ mask = assign_responsive(n, u, np.random.default_rng(0), "uniform", n_regions)
+ region = region_of(n, n_regions)
+ dead_per_region = [int((~mask[region == r]).sum()) for r in range(n_regions)]
+ assert all(0 < d < 100 for d in dead_per_region) # every region damaged, none wiped out
+
+
+def test_region_locality_keeps_peers_inside_the_region_and_stays_d_regular():
+ from pd.graph import build_graph, region_of
+ n, n_regions, degree = 2000, 10, 8
+ region = region_of(n, n_regions)
+ for locality, want in ((0.0, 0.1), (0.5, 0.5), (1.0, 1.0)):
+ cfg = SimConfig(n_nodes=n, degree=degree, n_regions=n_regions,
+ region_locality=locality, graph_seed=0)
+ g = build_graph(cfg)
+ assert np.all(np.diff(g.indptr) == degree) # exact d-regularity is preserved
+ same = float(np.mean(region[g.src] == region[g.indices]))
+ assert abs(same - want) < 0.05, (locality, same)
+
+
+def test_regional_churn_leaves_survivors_better_connected():
+ """The point of the correlated model: clustered failure removes whole neighbourhoods and
+ leaves the rest intact, so surviving nodes keep more live peers than under scattered failure."""
+ from pd.graph import build_graph
+ n, n_regions, degree, u = 4000, 20, 8, 0.4
+ cfg = SimConfig(n_nodes=n, degree=degree, n_regions=n_regions, region_locality=0.75,
+ graph_seed=0)
+ g = build_graph(cfg)
+ live_degree = {}
+ for mode in ("uniform", "regional"):
+ mask = assign_responsive(n, u, np.random.default_rng(1), mode, n_regions)
+ live_nbr = mask[g.indices] # is each peer alive?
+ counts = np.add.reduceat(live_nbr.astype(np.int32), g.indptr[:-1])
+ live_degree[mode] = float(counts[mask].mean()) # live peers of a surviving node
+ assert live_degree["regional"] > live_degree["uniform"] + 0.5