- include <memory> explicitly (std::unique_ptr; was transitive-only)
- validate env list elements: thread counts must be >0, payload sizes
and event payload >=0 — a negative size wrapped to a huge size_t in
makePayload, and atoi garbage became a 0-thread row
- emit -1 for p50/p99 in csv rows of lanes without latency samples
(async), so "not measured" is distinguishable from a measured 0 ns
- rename the scaling column to "vs first row" — the baseline is the
first configured thread count, not necessarily 1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measures the full C++ -> Nim -> C++ foreign round trip across setups:
scalar family (2 x int64 + float64), payload family (seq[byte], size
sweep incl. large payloads), and event delivery — through sync, async
(bounded in-flight window) and event lanes, over a thread-count sweep.
All handlers compute the same O(1) parity predicate, so setups differ
purely in transport cost. Every reply is verified; a bad reply or lost
event exits non-zero.
`nimble perf_cpp_e2e` regenerates the cpp bindings, builds libperfbench
with -d:danger over the NIM_FFI_MM matrix (orc + refc) — deliberately
bypassing the debug-orc nim_ffi_lib.cmake template build — and runs the
Release driver. NIM_FFI_PERF_* env knobs control threads, volume,
iterations and payload sizes; each table row also emits a csv line for
diff-friendly capture.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>