55 lines
1.7 KiB
Makefile
Raw Normal View History

VENV ?= .venv
PY := $(VENV)/bin/python
STAMP := $(VENV)/.installed
# Keep numpy/scipy BLAS single-threaded so joblib process parallelism doesn't oversubscribe.
export OMP_NUM_THREADS := 1
export OPENBLAS_NUM_THREADS := 1
export MKL_NUM_THREADS := 1
export NUMEXPR_NUM_THREADS := 1
pd: correlated AS/region churn, and two report caveats corrected Uncorrelated churn alone was incomplete: real outages take out a datacentre, AS or region as a unit. Adds failure domains and a correlated churn mode, plus the metric needed to tell the two apart. - n_regions / region_locality: nodes belong to equal-sized failure domains, and a configurable share of each node peers inside its own domain. Locality is what makes a failure domain a connectivity domain -- with region-blind peering, dropping whole regions removes a uniformly random set of nodes and is indistinguishable from uniform churn. The locality matchings keep the graph exactly d-regular (they change where peers are, never how many). - churn_mode = uniform | regional, swept per topology so both modes are compared on the same graph at an identical dead-node count. - frac_reached_live: coverage of the *responsive* network, alongside coverage of all nodes. The two move in opposite directions under correlated failure, so one number could not express the result. Measured (degree 4, 20 domains, 75% locality, half the network dead): clustered failure leaves the survivors fully connected -- live coverage 1.000 and delivery equal to the live-relay rate, i.e. nothing lost to routing -- where the same number of scattered failures gives 0.857 live coverage and loses delivery to broken routes. Correlated outages are gentler on the survivors than uniform churn, while stranding the dead domains. Verify check 8 anchors this. Also, per review of the caveats: exact d-regularity is a protocol requirement rather than a modelling simplification, and the timing-correlation adversary is deferred because it is only meaningful once the network emits cover traffic, which this simulator does not yet do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:37:56 +02:00
.PHONY: install smoke sweep sweep-fullscale redundancy percolation correlated-churn figures verify test lint clean
# The stamp is the real install; targets below depend on it so `make sweep` (etc.) auto-installs
# on a fresh checkout and re-installs whenever pyproject.toml changes.
$(STAMP): pyproject.toml
python3 -m venv $(VENV)
$(PY) -m pip install -U pip
$(PY) -m pip install -e ".[dev]"
@touch $(STAMP)
install: $(STAMP)
smoke: $(STAMP) ## fast end-to-end (seconds): tiny N, few rounds/seeds
$(PY) -m pd.sweep --config configs/smoke.yaml
sweep: $(STAMP)
$(PY) -m pd.sweep --config configs/default.yaml
sweep-fullscale: $(STAMP)
$(PY) -m pd.sweep --config configs/fullscale.yaml
Add linkability, messaging redundancy and churn percolation to pd; report Extends the pd Blend simulator along two axes the deanonymization model opened up, adds the reports/blend/pd report of record, and fixes three correctness defects found while reviewing the result. Linkability over time (pd.linkability): - time to link an emitter ~ 30s*ln(1/(1-alpha))/(stake*q): inversely proportional to stake, so a 5% staker is linked in ~2 days and a 0.001% staker only after ~27 years; - time to certify a node's stake >= theta from the count of attributable observations (relative precision ~1/sqrt(N)): sizing a node costs 100-400x more than identifying it, and sub-0.1% stake is practically unlearnable. Both are closed forms over the exact deanonymization rates and a stake-proportional 30 s emission cadence, checked against a Monte-Carlo of the emission process in verify. Messaging redundancy (R independent cascades per emission, R = 1..4): - `redundancy` knob threaded through config/rng/propagation/engine/metrics/ sweep; a node receives from whichever cascade reaches it first, so arrival times combine element-wise. Delivery and capture both follow 1-(1-x)^R, so redundancy trades reliability against anonymity and divides time-to-link by ~R. Measured: delivery 0.34 -> 0.81 at 30% churn for R = 1 -> 4, while a 1%-staker's time to link falls 10 d -> 2.5 d. - Redundancy buys NO coverage: a cascade only delivers if the sender could already route to its relay, so every delivered cascade floods the sender's own component. Coverage is flat in R to four decimals at every degree. - Near the percolation threshold the cascades fail together rather than independently, so redundancy under-delivers against 1-(1-p1)^R there. Churn percolation (configs/percolation.yaml, verify check 7): - the flood only crosses responsive nodes, so it lives on the responsive sub-graph -- site percolation on a d-regular graph. A network survives churn only up to u_c = 1 - 1/(degree-1); measured collapse lands on the predicted threshold for every degree (3 -> 0.50, 6 -> 0.80, 16 -> 0.93), which inverts into the sizing rule degree > 1 + 1/(1-u). Correctness fixes: - redundancy delay used the fastest cascade's own full delay, which over-states it (min-max vs max-min); now the element-wise earliest arrival, reducing exactly to the single-cascade model at R = 1 (test); - the "redundancy improves coverage" claim was false in both the report and the simulator README -- removed and replaced with the measured result; - per-hop latency is degree-dependent (1.5 s at degree 16 to 2.7 s at degree 3), not a flat 1.6 s; and the worst-case observation figure was averaged over degrees -- at degree 8 and f_adv = 0.2 it is 0.83 -> 1.000. Statistics: round counts raised for resolution rather than speed -- 8000 rounds per cell in the main sweep, 9600 in the redundancy study, 6400 in the percolation study, giving SEM <= 0.009 on every delivery rate and <= 0.04 s on every delay mean. The previous redundancy grid (144 rounds/cell) produced a non-monotonic delivery curve; it is now monotonic and within 0.015 of theory. Adversary and deanonymization metrics remain closed-form and exact. reports/blend/pd: the report of record -- peering-degree trade-offs across speed, observation, eclipse, deanonymization and reliability, plus the time-to-link, stake-inference, redundancy and churn-threshold sections, with 21 figures of record and an explicit sampling-error statement. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 22:37:22 +02:00
redundancy: $(STAMP) ## messaging redundancy R=1..4 (delivery vs deanonymization, time-to-link)
$(PY) -m pd.sweep --config configs/redundancy.yaml
percolation: $(STAMP) ## churn threshold: coverage collapse at u_c = 1 - 1/(degree-1)
$(PY) -m pd.sweep --config configs/percolation.yaml
pd: correlated AS/region churn, and two report caveats corrected Uncorrelated churn alone was incomplete: real outages take out a datacentre, AS or region as a unit. Adds failure domains and a correlated churn mode, plus the metric needed to tell the two apart. - n_regions / region_locality: nodes belong to equal-sized failure domains, and a configurable share of each node peers inside its own domain. Locality is what makes a failure domain a connectivity domain -- with region-blind peering, dropping whole regions removes a uniformly random set of nodes and is indistinguishable from uniform churn. The locality matchings keep the graph exactly d-regular (they change where peers are, never how many). - churn_mode = uniform | regional, swept per topology so both modes are compared on the same graph at an identical dead-node count. - frac_reached_live: coverage of the *responsive* network, alongside coverage of all nodes. The two move in opposite directions under correlated failure, so one number could not express the result. Measured (degree 4, 20 domains, 75% locality, half the network dead): clustered failure leaves the survivors fully connected -- live coverage 1.000 and delivery equal to the live-relay rate, i.e. nothing lost to routing -- where the same number of scattered failures gives 0.857 live coverage and loses delivery to broken routes. Correlated outages are gentler on the survivors than uniform churn, while stranding the dead domains. Verify check 8 anchors this. Also, per review of the caveats: exact d-regularity is a protocol requirement rather than a modelling simplification, and the timing-correlation adversary is deferred because it is only meaningful once the network emits cover traffic, which this simulator does not yet do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 11:37:56 +02:00
correlated-churn: $(STAMP) ## correlated AS/region outages vs uniform churn, matched fractions
$(PY) -m pd.sweep --config configs/correlated-churn.yaml
figures: $(STAMP) ## make figures RUN=runs/<dir>
$(PY) -m pd.plotting.make_figures --run $(RUN)
verify: $(STAMP)
$(PY) -m pd.verify
test: $(STAMP)
$(PY) -m pytest
lint: $(STAMP)
$(VENV)/bin/ruff check src scripts tests
clean:
rm -rf runs/* figures/* .pytest_cache .ruff_cache .mypy_cache