logos-messaging-nim/tests/messaging/test_rln_proof_attach.nim

117 lines
4.1 KiB
Nim
Raw Normal View History

feat: attach and refresh RLN proofs in the send service (Stream B) (#4069) * feat: attach RLN proofs at the SendService transmission stage The relay send path published without an RLN proof: proof generation lived client-side in (legacy)lightpushPublish, so messages dispatched through SendService -> RelaySendProcessor reached the network unproven and would be rejected by an RLN-enforcing relay. Adds Waku.attachRlnProof in the waku/api publish surface and calls it from SendService immediately after admission, in both send() and the retry loop. Placement is load-bearing: - After admit(), so a message rejected by the rate limiter never draws a nonce. - At transmission rather than API entry, because a proof binds to the epoch current when the message goes out, and a task can be retried for up to MaxTimeInCache after send() returns. attachRlnProof is a no-op without RLN mounted (message passes through unproven, as today) and short-circuits on a message that already carries a proof, so retrying a task neither redraws a nonce nor changes the bytes. It uses generateRLNProofWithRootRefresh rather than the plain generator: a task can wait in the task cache while the group root moves on chain, so the proof is validated against the acceptable-root window and regenerated once against a refetched merkle path if it went stale. Proof-generation failure parks the task as NextRoundRetry rather than failing it, matching the admission path: the dominant failure is NonceLimitReached (RLN's own per-epoch budget exhausted), which the service loop resolves as the epoch rolls over. Adds tests/messaging/test_rln_proof_attach.nim covering the unmounted pass-through, attach when mounted, and the idempotency contract that the retry loop depends on. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: handle RLN publish rejections in the send service retry loop An RLN-invalid publish rejection now recovers through the send service's existing retry loop instead of an inline retry in the kernel. When a relay or lightpush publish is rejected as RLN-invalid, the processor clears the message's stale proof, schedules a background merkle-proof refresh, and parks the task as NextRoundRetry. The next loop round re-admits the task and regenerates the proof against the refreshed path. Clearing the proof is required: attachRlnProof short-circuits on a message that already carries one, so without the clear the task would resend the rejected proof until it ages out. The relay processor previously failed such tasks outright, with no recovery. Kernel changes supporting this: - Remove runRlnRefreshRetry from legacyLightpushPublish. The legacy path now schedules the refresh and returns the error tagged with RlnProofRefreshScheduledMsg, matching the non-legacy path; retrying is the caller's decision. Drops the now-unused RlnMerkleProofRefreshTimeout. - generateRLNProofWithRootRefresh reuses the nonce drawn for the first attempt when it regenerates after a stale root, rather than drawing a second. Only the merkle path differs between the two attempts, so a redraw would spend two message ids from the epoch budget on a single message and drift the rate limit manager's accounting away from the nonce manager's. Adds Waku.isRlnRejection / Waku.onRlnProofRejected as the messaging layer's handle on the kernel's RLN rejection detection and background refresh. Updates the legacy lightpush tests to the schedule-refresh contract. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: cover currentRlnEpochQuota (mounted and unmounted) currentRlnEpochQuota ships in the enforcement PR, but its mounted-RLN assertion needs the anvil-backed group-manager scaffolding that lives in this file, so the coverage rides along here: none when RLN is unmounted, and the epoch index + userMessageLimit when it is. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor: consolidate the RLN-rejection parking into parkForRlnProofRefresh The relay and lightpush processors duplicated the RLN-rejection recovery (schedule background refresh, clear the stale proof, reset admission, park as NextRoundRetry); both now call parkForRlnProofRefresh in send_processor, so the proof-clear the retry contract depends on cannot drift between the two processors. Also resets firstAdmittedTime so the regenerated proof's fresh nonce is re-admitted rather than sent uncharged. No behavior change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: clarify that lightpush reuses an already-attached RLN proof The old wording ("attaches an RLN proof per attempt") reads as if every retry redraws a nonce. The flow proves a message only when it carries no proof, so a task admitted once reuses its proof and nonce across retries. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: retry once on scheduled RLN proof refresh in the lightpush REST handlers The kernel lightpush publish paths no longer retry an RLN-invalid publish inline. On a stale merkle root they schedule a background cache refresh and return early, tagging the error with RlnProofRefreshScheduledMsg — retrying is the caller's decision, so the send service recovers through its own loop and the kernel exposes mechanism only. The synchronous REST endpoints have no such loop: they call publish once and map the result to an HTTP status. Left unchanged, a transient stale-root rejection that the kernel previously absorbed via runRlnRefreshRetry would now surface to the HTTP client as a 503. Restore the transparent retry where it belongs under this layering — at the caller — instead of back in the kernel where it would re-nest inside the send service's retry. Both the legacy and v3 handlers now retry the publish exactly once when the first result carries RlnProofRefreshScheduledMsg, under the same FutTimeoutForPushRequestProcessing bound. The handler's message carries no proof, so the retry regenerates against the refreshed merkle path. Any other error, and any error on the retry itself, maps to its response as before. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: retry RLN proof attach on a charged-but-unproven task admitAndProve set firstAdmittedTime before attaching the proof, then guarded its whole body on firstAdmittedTime.isSome(). A transient proof attach failure (e.g. NonceLimitReached) left the task charged but with an empty proof, and the next round's early-return skipped the attach entirely and shipped the message bare. The docstring's own invariant — "once admitted, a task keeps its slot and its proof" — was violated: it kept the slot but not the proof. Guard only the rate-limit charge on firstAdmittedTime, not the attach. attachRlnProof is already idempotent (short-circuits when RLN is unmounted or a proof is present), so it is safe to call every round: a charged-but- unproven task retries the attach until it sticks, then short-circuits. The ordering invariant holds (charge strictly before attach, so an over-budget message never draws a nonce), the charge stays once-per-task, and NO_PEERS retries remain free. This also removes a latent relay double-charge: previously a bare message reached the relay, was rejected as RLN-invalid, and parkForRlnProofRefresh reset firstAdmittedTime — re-charging a slot on the next round. The message now never leaves unproven. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 10:15:43 +02:00
{.used.}
import std/[options, net, osproc]
import chronos, testutils/unittests, results, stew/byteutils
import
logos_delivery/waku/[waku, waku_core, rln],
logos_delivery/waku/node/waku_node,
logos_delivery/waku/node/waku_node/relay,
logos_delivery/waku/api/publish,
logos_delivery/api/conf/messaging_conf,
logos_delivery/waku/factory/waku_conf
import
../testlib/testasync,
../waku_rln_relay/utils_onchain,
../waku_rln_relay/rln/waku_rln_relay_utils
proc testConf(): WakuConf =
var conf = MessagingClientConf()
.toWakuNodeConf(messaging_conf.LogosDeliveryMode.Core).valueOr:
raiseAssert error
conf.listenAddress = parseIpAddress("0.0.0.0")
conf.tcpPort = Port(0)
conf.discv5UdpPort = Port(0)
conf.clusterId = Opt.some(3'u16)
conf.numShardsInNetwork = 1
conf.rest = false
return conf.toWakuConf().valueOr:
raiseAssert error
proc testMessage(): WakuMessage =
WakuMessage(
payload: "hello".toBytes(),
contentTopic: "/test/1/attach/proto",
timestamp: 1_700_000_000_000_000_000,
)
suite "SendService RLN proof attach":
asyncTest "passes the message through unproven when RLN is not mounted":
## The default (no-RLN) configuration must be unaffected: no proof is
## attached and the message reaches the send processors unchanged.
let waku = (await Waku.new(testConf())).expect("Waku.new")
let msg = testMessage()
let attached = (await waku.attachRlnProof(msg)).expect("attachRlnProof")
check:
attached.proof.len == 0
attached.payload == msg.payload
attached.contentTopic == msg.contentTopic
asyncTest "currentRlnEpochQuota is none when RLN is not mounted":
## The rate limit manager reads `none` as "use the wall-clock fallback".
let waku = (await Waku.new(testConf())).expect("Waku.new")
check waku.currentRlnEpochQuota().isNone()
suite "SendService RLN proof attach - RLN mounted":
var
waku {.threadvar.}: Waku
anvilProc {.threadvar.}: Process
manager {.threadvar.}: OnchainGroupManager
asyncSetup:
anvilProc = runAnvil(stateFile = Opt.some(DEFAULT_ANVIL_STATE_PATH))
manager = waitFor setupOnchainGroupManager(deployContracts = false)
waku = (await Waku.new(testConf())).expect("Waku.new")
await waku.node.setRlnValidator(
getWakuRlnConfig(
manager = manager,
userMessageLimit = 20,
index = MembershipIndex(1),
epochSizeSec = 600,
)
)
let credentials = generateCredentials()
(
waitFor cast[OnchainGroupManager](waku.node.rln.groupManager).register(
credentials, UserMessageLimit(20)
)
).isOkOr:
assert false, "failed to register RLN credentials: " & error
asyncTeardown:
## The RLN proof-generator provider is registered on the global broker
## context; without stopping RLN it leaks into the next test's setup.
try:
await waku.node.rln.stop()
except Exception:
assert false, "failed to stop RLN: " & getCurrentExceptionMsg()
stopAnvil(anvilProc)
asyncTest "attaches a proof":
let attached = (await waku.attachRlnProof(testMessage())).expect("attachRlnProof")
check attached.proof.len > 0
asyncTest "currentRlnEpochQuota reports RLN's epoch and user message limit":
## Wires the rate limit manager to RLN: the manager clamps its configured
## cap to `messageLimit` and rolls on `epochIndex`.
let quota = waku.currentRlnEpochQuota()
check:
quota.isSome()
quota.get().messageLimit == 20'u64 # the mounted userMessageLimit
quota.get().epochIndex > 0'u64 # unixTime div epochSize, far from zero
asyncTest "is idempotent: a message that already carries a proof is untouched":
## Pins the retry contract: the send service re-attaches on every round, so
## re-attaching must neither draw a fresh nonce nor change the bytes —
## otherwise a retried task would resend under a new nullifier.
let first = (await waku.attachRlnProof(testMessage())).expect("first attach")
let second = (await waku.attachRlnProof(first)).expect("second attach")
check:
first.proof.len > 0
second.proof == first.proof