logos-messaging-nim/apps/networkmonitor/networkmonitor_metrics.nim

108 lines
3.4 KiB
Nim
Raw Normal View History

{.push raises: [].}
import
std/[net, json, tables, sequtils],
chronicles,
chronicles/topics_registry,
chronos,
json_serialization,
metrics,
metrics/chronos_httpserver,
presto/route,
presto/server,
results
logScope:
topics = "networkmonitor_metrics"
# On top of our custom metrics, the following are reused from nim-eth
#routing_table_nodes{state=""}
#routing_table_nodes{state="seen"}
#discovery_message_requests_outgoing_total{response=""}
#discovery_message_requests_outgoing_total{response="no_response"}
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_type_as_per_enr,
"Number of peers supporting each capability according to the ENR",
labels = ["capability"]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_cluster_as_per_enr,
"Number of peers on each cluster according to the ENR", labels = ["cluster"]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_type_as_per_protocol,
"Number of peers supporting each protocol, after a successful connection) ",
labels = ["protocols"]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_user_agents,
"Number of peers with each user agent", labels = ["user_agent"]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicHistogram logos_delivery_networkmonitor_peer_ping,
"Histogram tracking ping durations for discovered peers",
buckets = [10.0, 20.0, 50.0, 100.0, 200.0, 300.0, 500.0, 800.0, 1000.0, 2000.0, Inf]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_count,
"Number of discovered peers", labels = ["connected"]
refactor(metrics): give every metric a logos_delivery_ prefix (#4074) * refactor(metrics): prefix node metrics with logos_delivery_ Every metric the node exports now starts with logos_delivery_. There was no prefix mechanism before: nim-metrics derives the exported name from the Nim identifier, no declaration passed an explicit `name = "..."`, and the waku_ convention was maintained by hand -- 69 of the 80 node metrics followed it and 11 did not (query_count, query_time_secs, event_loop_load, event_loop_accumulated_lag_secs, postgres_payload_size_bytes, reconciliation_*, total_* and the camelCase rendezvousPeerFoundTotal). Identifiers are renamed rather than given a `name = "..."` argument, keeping the invariant that the Nim identifier is the exported name and letting the compiler check every call site. rendezvousPeerFoundTotal becomes logos_delivery_rendezvous_peer_found: it was the only camelCase metric, and the trailing Total was redundant since nim-metrics already appends _total to counters at exposition time. library/ is untouched on purpose -- `proc waku_version()` in kernel_api/debug_node_api.nim is the exported libwaku C ABI symbol, not the gauge of the same name in node_telemetry.nim. BREAKING CHANGE: metric names change. Dashboards, alert rules and recording rules that reference waku_* must be updated; see docs/operators/how-to/monitor.md. * refactor(metrics): prefix auxiliary app metrics with logos_delivery_ Applies the same prefix to the tools shipped from this repo: liteprotocoltester (lpt_*), networkmonitor (networkmonitor_*), chat2bridge (chat2_*) and the lightpush_mix example (lp_mix_*). These tools are not the delivery node and already had their own consistent prefixes, so this commit is separable from the node rename if the intent was to namespace only the node itself. * chore(metrics): query old and new metric names in Grafana dashboards 212 expressions across 9 dashboards now match both the waku_* and the logos_delivery_* spelling, so panels keep working across the upgrade and over historical data: sum by (type)((increase(waku_node_errors_total{...}[$__rate_interval]) or increase(logos_delivery_node_errors_total{...}[$__rate_interval]))) The `or` is placed around the leaf, inside every aggregation. That depth is load-bearing: `or` keeps its right operand only for label sets absent from the left, so `sum by (type)(old) or sum by (type)(new)` aggregates each half of the fleet separately and then discards the right one entirely -- silently dropping every already-upgraded node. 36 panels here collapse `instance`. Measured against a local Prometheus scraping two targets, one exporting old names at 10/s and one exporting new names at 20/s (truth 30/s): union outside the aggregation gives 10, union around the leaf gives 30. Where the leaf sits in a range vector the whole call is duplicated, since `(a or b)[5m]` is not valid PromQL. Once every scraped node runs a release with the new names and the old samples have aged out of retention, the `or` half can be deleted. * test(e2e): expect logos_delivery_-prefixed metric names The e2e suite asserts against a live /metrics endpoint, which serves only the new names, so these are replaced rather than unioned. libp2p_* entries are unchanged. * docs(operators): document the logos_delivery_ metric prefix Records that every metric the node exports is prefixed, that dependency metrics (libp2p_*, nim_gc_*, process_*) keep their own names, and shows where the `or` has to sit if operators maintain their own dashboards or alert rules. * refactor(metrics): name the store fleet metrics after store, not relay logos_delivery_relay_fleet_store_msg_size_bytes and _msg_count are declared in waku_store/protocol_metrics.nim and recorded by the store client, but carried a relay prefix. Renamed to logos_delivery_store_fleet_msg_size_bytes and logos_delivery_store_fleet_msg_count. The dashboard keeps matching the old exported name, which was waku_relay_fleet_store_*. Note that both metrics are wrong independently of their name, see the PR description.
2026-07-29 18:56:50 +01:00
declarePublicGauge logos_delivery_networkmonitor_peer_country_count,
"Number of peers per country", labels = ["country"]
type
CustomPeerInfo* = object # populated after discovery
lastTimeDiscovered*: int64
discovered*: int64
peerId*: string
enr*: string
ip*: string
enrCapabilities*: seq[string]
country*: string
city*: string
maddrs*: seq[string]
# only after ok connection
lastTimeConnected*: int64
retries*: int64
supportedProtocols*: seq[string]
userAgent*: string
lastPingDuration*: Duration
avgPingDuration*: Duration
# only after a ok/nok connection
connError*: string
CustomPeerInfoRef* = ref CustomPeerInfo
# Stores information about all discovered/connected peers
CustomPeersTableRef* = TableRef[string, CustomPeerInfoRef]
# stores the content topic and the count of rx messages
ContentTopicMessageTableRef* = TableRef[string, int]
proc installHandler*(
router: var RestRouter,
allPeers: CustomPeersTableRef,
numMessagesPerContentTopic: ContentTopicMessageTableRef,
) =
router.api(MethodGet, "/allpeersinfo") do() -> RestApiResponse:
let values = toSeq(allPeers.values())
return RestApiResponse.response(values.toJson(), contentType = "application/json")
router.api(MethodGet, "/contenttopics") do() -> RestApiResponse:
# TODO: toJson() includes the hash
return RestApiResponse.response(
$(%numMessagesPerContentTopic), contentType = "application/json"
)
proc startMetricsServer*(serverIp: IpAddress, serverPort: Port): Result[void, string] =
info "Starting metrics HTTP server", serverIp, serverPort
try:
startMetricsHttpServer($serverIp, serverPort)
except Exception as e:
error(
"Failed to start metrics HTTP server",
serverIp = serverIp,
serverPort = serverPort,
msg = e.msg,
)
info "Metrics HTTP server started", serverIp, serverPort
ok()