mirror of
https://github.com/logos-co/logos-logoscore-cli.git
synced 2026-08-30 20:31:09 +00:00
master
14
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e9f134f5aa |
fix: draw the line at LOADED, not merely known — for both watch and call (#109)
* fix(watch): defer the subscription instead of refusing a module that is not up yet
`logoscore watch <module>` answered WATCH_FAILED and exited whenever the
module's registry socket had no listener at that instant, and nothing ever
retried. Two ordinary states hit it:
* subscribing before the module is loaded at all;
* subscribing in the window after `load-module` RETURNS but before the
module publishes its object. `load-module` answers once the plugin is
in; publishing happens afterwards.
Measured on macOS, cold first load of a freshly built plugin:
.803 Registry connect attempt started (no peer contact yet)
.804 LogosAPIConsumer: Requesting object: "modules_state"
.804 Warning: Not connected to registry. Cannot request object
.957 [modules_state] RemoteTransportHost: Published object
Refused at .804; published 153 ms later. The refusal is
LogosAPIConsumer::requestObject's `isConnected()` guard --
`m_connected && endpointHasListener()`, a socket liveness probe. This is the
third un-migrated instance of the one-shot subscribe shape; the generated Qt
wrappers and the QML bridge both moved to onEventWhenAvailable already.
The dangerous part was never the failure, it was that nothing checked it: a
script's only event assertion went dead while every later line still passed.
It is also inconsistent with the call path, for identical input. Same module,
same instant, not loaded:
watch -> WATCH_FAILED in 139 ms
call -> RPC_FAILED in 20 544 ms
invokeRemoteMethod reaches the transport through acquireCachedObject and never
sees the guard, so it waits out waitForSource(20s). One path gave up in a
millisecond, the other waited 20 seconds.
Named events now go through onEventWhenAvailable, which holds the subscription,
arms it when the module appears (including one loaded later in the session) and
re-arms across a reconnect. The wildcard form (`watch <module>` with no
--event) cannot: it subscribes with an EMPTY event name, which LogosObject
reads as "every event" but onEventWhenAvailable refuses outright -- so a fix
covering only the named form would have left the CLI's DEFAULT invocation
silently dead. It goes through whenObjectAvailable instead.
Deferring must not swallow a typo, so a name the host has never heard of still
fails fast rather than parking on a subscription with no future. Before this
change both cases failed identically and that was free; now it is pinned.
Three tests in each integration suite, against a fixture whose module is
discoverable but NOT loaded -- deterministic, where reproducing via a
post-load-module subscribe needs a cold machine. Verified as a negative
control: with the tests present and this fix reverted, both "arms later" tests
fail with the exact production signature,
{"code":"WATCH_FAILED","message":"Failed to watch events for module
'test_basic_module'.","status":"error"}
and WatchUnknownModuleStillFailsFast passes, confirming it pins existing
behaviour rather than the new path. With the fix: checks.tests-logoscore and
checks.tests-logosctl both green (250 + 30 + 27 tests), the two arming tests
dropping from 64s/70s of failing poll loop to 1.8s/1.9s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(watch): draw the line at LOADED, not merely known
Follow-up to the previous commit, which gated on getKnownModuleNames() and so
deferred a subscription for any module the host had ever heard of. That made
`watch` on an unloaded module park forever on a subscription that might never
arm — measured here as `timeout 12` reporting 124, with no answer printed at
all. A module that is not loaded may never be, so refusing is the useful
answer, and a typo now gets that same answer rather than a worse one.
The contract is two halves:
* NOT LOADED -> fail fast. Already true before any of this; now pinned.
* LOADED -> always succeed. This is the half that was broken, and the
reason the deferred path is here at all: `load-module` returns once the
plugin is in, but the module publishes its object afterwards, so there is a
window where the host reports a module loaded and the one-shot
requestObject() refused it. Cold, that window is seconds.
Past the loaded check the subscription cannot fail, by construction:
onEventWhenAvailable / whenObjectAvailable answer 0 only for arguments they
refuse (empty name, null callback), none reachable here, and neither touches
the transport on the calling thread.
COVERAGE, stated plainly because it got weaker and the weakness is not visible
from a green run. The unloaded half is deterministic — the fixture's daemon
starts with the module discoverable but not loaded, and that state holds still.
The loaded half is NOT, and cannot be made so from the CLI: nothing can hold a
module in "loaded but not yet published". Verified rather than assumed — with
these tests against the ORIGINAL master implementation, all three PASS on a warm
machine. They bite on a cold CI worker, which is where the failure was seen,
but they are not a warm-machine regression net.
Two deterministic alternatives were considered and rejected. Subscribing before
the load reproduced it reliably, and was what the previous commit tested, but it
is no longer valid behaviour under this contract. Crashing the module leaves it
briefly listed as loaded with its socket gone — the right state, reached by a
race, since the daemon's SIGCHLD handler flips it out of "loaded"
asynchronously. A racing test is worse than none.
checks.tests-logoscore and checks.tests-logosctl both green (250 + 30 + 27).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(call): draw the line at LOADED, not merely known
The same line the previous commit drew for `watch`. `callModuleMethod` had no
precondition at all: its `!moduleClient` guard can never fire, because
LogosAPI::getClient constructs a client for any name and never returns null.
So an absent module was found the expensive way, and two 20s deadlines raced to
report it — the daemon's acquire (object_unavailable) against the CLI client's
own RPC deadline (RPC_FAILED, a client-side literal the daemon cannot emit).
Measured: 17 calls split 5/12 across the two codes, all at 19.99-20.43s.
Not fixed by a shorter deadline: one Timeout feeds both the acquire and the
call, and the long acquire is the documented startup contract. The daemon
already has an in-process answer and already uses it in watchModuleEvents.
core_service is exempt (published by the daemon's own provider, never in the
loaded set). The startup window is unaffected: liblogos marks loaded before
publish, so a warming module passes the gate and keeps its full budget.
Tests assert the property, not a tally — the old behaviour was bimodal per run,
so sampling can pass unfixed. A positive control covers the load->publish
window, which is what a shortened deadline would break. Two assertions in
NoLoadNegativePaths were passing BY hanging (timeout's 124 is non-zero); both
now exclude it.
Follow-ups, not in this commit: logos-test-modules conformance still expects
["object_unavailable","RPC_FAILED"] for failure/A/module-not-loaded and must
move with the relock; test_integration_logoscore.cpp has the same two
hang-masked assertions.
checks.tests green: logosctl 250 + 30 + 29, logoscore 20 + 27.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(deps): protocol 303ab08 — the last level of the carrier chain
Bumps logos-protocol, logos-plugin-qt and logos-liblogos together. The three
levels below are logos-plugin-qt#30, logos-qt-sdk#46 and logos-liblogos#190.
Bumping ONLY this repo's own logos-protocol input does not work, and that is not
a lock-tidiness opinion -- it was measured. That bump moves 1 of 230 protocol
nodes, builds green, and ships a liblogos_protocol.dylib WITHOUT the change,
because the runtime library is staged by a CARRIER rather than by this input. Of
everything in the closure only logos-protocol, logos-plugin-qt and
logos-liblogos carry one; cpp-sdk, capability-module, package-manager,
package-downloader and test-modules do not. Nothing fails when you get this
wrong -- the lock diff is real, the build is green, and the daemon loads a
library without the fix.
The lock still holds 230 protocol nodes at 16 revs afterwards. That is the usual
explosion and it is not what ships; the closure is the claim:
protocol paths in the built daemon closure: exactly ONE
n7yilrs2...-logos-protocol-lib-0.8.0 (the build of 303ab08)
shipped liblogos_protocol.dylib:
"(any)" (UTF-16) x1 <- present only in 303ab08
old warning x0
new warning x2 <- whenObjectAvailable + the fixed onEventWhenAvailable
Both halves of that are needed. `strings` cannot see "(any)" because
QStringLiteral is UTF-16, and the new warning text appears once even in the OLD
library because whenObjectAvailable has always used it -- checking either alone
reports the wrong answer, which it did here first time round.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(watch): both subscription forms are now one call
logos-protocol#74 let onEventWhenAvailable take an EMPTY event name as the
wildcard, so the detour the wildcard form needed is gone: whenObjectAvailable()
+ requestObject() + onEvent(), a QPointer guard and a nested lambda collapse
into the same one-line call the named form already made.
The detour existed because ONE guard rejected three unrelated arguments at once
-- an empty object name and a null callback, which are unusable, and an empty
event name, which is meaningful and which the plain onEvent has always honoured.
Routing around it kept `watch <module>` with no --event working, but left the
CLI's DEFAULT invocation on a different code path from its --event form, which
is the shape a silent regression hides in.
Requires the lock bump in the previous commit: against the old library
onEventWhenAvailable answers 0 for an empty name, so the wildcard would refuse.
That is not a latent trap -- WildcardWatchSucceedsTheInstantLoadModuleReturns
fails loudly on exactly it, in both suites.
checks.tests-logoscore and checks.tests-logosctl both green (20+27 and
250+30+29), status read from nix rather than from a pipeline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(call): mirror the loaded-set gate cases into the logoscore suite
The gate landed with coverage in tests/test_integration.cpp only, so the
`logoscore` binary had none for it -- and the two binaries compile from separate
files, so "the other suite covers it" is not true here. The `watch` cases for
the same gate are already in both.
Ports both cases plus the hardening of NoLoadNegativePaths: `timeout` also exits
non-zero, so its two "call must not succeed" assertions passed BY hanging, and
now exclude 124 explicitly.
Verified as coverage, not as compilation. With the gate deleted from
callModuleMethod and nothing else changed:
FAILED ErrorPathTest.NoLoadNegativePaths (25012 ms)
FAILED ErrorPathTest.CallOnAModuleThatIsNotLoadedIsRefusedNotAwaited (10929 ms)
OK ErrorPathTest.CallImmediatelyAfterLoadStillReachesTheModule (924 ms)
The two negative cases fail on the timeouts they exist to forbid, and the
positive control stays green -- which is the point of shipping it alongside
them: it is what a shortened acquire deadline would break while the negative
cases still passed. Restored, both suites green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
500f5de1f1 |
chore: bump protocol to 0.9 (#105)
* chore: bump protocol to 0.9 * chore(deps): relock liblogos host stack * bump * bump * test(access-policy): use token-bound callers |
||
|
|
162dbc9fff |
fix(daemon stop): stop losing the shutdown reply, and stop calling that a failure
`logosctl daemon stop` printed {"code":"RPC_FAILED","message":"shutdown RPC
call failed."} and exited 3 for shutdowns that had already succeeded. It cost
the "Stop the daemon" step of doctests/logosctl-daemon.test.yaml one failure
out of nine identical shutdowns in the same macOS CI job; the daemon really
had stopped, and `daemon status` two seconds later said so.
Two independent defects, one on each side of the call.
DAEMON. CoreServiceImpl::shutdown() returned {"status":"ok"} and left the
event loop from a detached std::thread that slept 200ms and called
QCoreApplication::quit(). The reply is not on the wire at that point: the
transport serialises it after the handler returns and hands it to the socket,
which only pushes it out when the event loop services that socket's write
notifier. quit() is not a queued event -- QCoreApplication::exit() interrupts
the dispatcher directly -- so if the main thread was descheduled for longer
than the sleep, the loop came back, exited, and the buffered reply died with
the process. QtRO surfaces no transport error for this; the client just waited
out its 20s deadline and saw nothing.
The quit now runs on the main thread, from a timer, and drains the event loop
before ending it. The 200ms is now a courtesy margin rather than the
correctness mechanism, and $LOGOSCTL_SHUTDOWN_GRACE_MS makes it settable --
including to 0, which the new regression test uses because it is the setting
that used to lose the reply outright.
QtRO offers nothing better: QRemoteObjectHostBase has no per-reply
write-completion signal and no client-disconnect signal, so "quit when the
response has actually been flushed" is not reachable without forking Qt, and
the daemon also serves plain TCP/TLS through a different transport.
CLIENT. RpcClient::shutdown() reported RPC_FAILED whenever the reply was not
an object -- including when there was no reply. But a missing reply is the
expected outcome of asking a process to die, and both docs said so already:
docs/spec.md promised "the client treats the connection loss as a successful
shutdown" and docs/project.md promised exit 0 for it. Neither was implemented.
It now answers the question the reply was standing in for, from evidence: the
pid recorded in daemon/state.json (snapshotted before the call, since a clean
shutdown deletes that file) is watched for up to 15s, or for a remote daemon
the endpoint is re-probed. Gone means success, with `confirmed_by` naming the
evidence; still running means a real error, with a message that says which.
Blindly treating silence as success would have been the more dangerous
mistake -- a wedged daemon is also silent -- so it is not what this does.
That inference is only sound about a pid that was alive to begin with, so
`stop` now refuses a stale session up front the way `daemon status` already
does: a state.json naming this client's instance and a dead pid means there is
no daemon to stop (NO_DAEMON, exit 2). Without it, a session left behind by
last week's daemon would "connect" to nothing, time out, observe that the pid
is gone, and call that a successful shutdown.
TESTS.
* ShutdownReplyTest.StopSucceedsWithNoGracePeriod (integration): 60
start/stop cycles at LOGOSCTL_SHUTDOWN_GRACE_MS=0, asserting the command
succeeds, the daemon is actually gone, and the reply arrived rather than
being reconstructed from the process exiting. Measured through this
fixture on macOS: 6 losses in 100 cycles before the daemon fix, 0 in 120
after.
* CommandTest.Stop_StaleSession_* : the stale-session guard, its live-pid
control, and the remote-client case it must not block. CommandTest now
isolates HOME and the config dir, so the suite no longer reads whichever
~/.logosctl the developer happens to have.
* ProcessUtil.WaitForProcessExit* : the primitive the confirmation rests on.
A/B over the shipped binaries, 30 stop cycles per arm at zero grace, macOS:
pre-fix 8 failures; daemon fix only 0 (no reply lost); client fix only 0
(20 replies lost, every command still correct); both 0. At the default 200ms
grace both arms are clean, which is why this presented as a rare CI flake.
Independent of PR #99: that PR does not touch either function, and the two
diffs do not overlap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ed19258375 |
fix(core_service): report METHOD_FAILED from the error channel, not a null value
callModuleMethod judged failure with `ret.is_null()` because it called the
one invokeRemoteMethod overload that has no CallError* parameter
(logos_api_client.h:352-356, whose body forwards to the QVariant overload
with the error channel dropped). A method that legitimately returns null was
therefore indistinguishable from a call that failed — and that single line
was the entire empirical basis for the qt-generator's refusal to allow an
optional return.
Switching to the CallError-carrying overload (logos_api_client.h:98) is not
sufficient on its own: an unknown method name is deliberately NOT reportable
on the wire (logos_protocol.h:274-279 says so outright, and the cdylib
dispatch ends `return nullptr; // unknown method`), so a naive !err.ok()
would have turned every typo into a silent success. The decision is now:
!err.ok() -> METHOD_FAILED + {code,message,origin}
result is a dispatch_failed envelope -> METHOD_FAILED (the provider refused)
null AND method provably not exposed -> METHOD_NOT_FOUND + available_methods
otherwise -> ok, null included
METHOD_NOT_FOUND is not invented — docs/spec.md:918 specified that envelope,
with available_methods, all along; core_service simply never produced it. It
costs one extra round-trip only on a null return.
The logic lives in a new pure unit, core_service/call_envelope.{h,cpp}, with
no Qt and no logos-protocol, which is what makes it unit-testable at all. The
value path is byte-identical: the same nlohmannArgsToQVariantList /
qvariantToNlohmann the json overload used internally.
Behaviour changes a reviewer must agree with: a null return is now `ok`
rather than METHOD_FAILED, and a dispatch_failed envelope returned as data is
now METHOD_FAILED rather than `ok`. No existing test encoded the old
behaviour; no exit code or ok/error verdict flipped in any fixture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0b2afed18a |
chore(deps): track master for protocol, cpp-sdk and plugin-qt (#92)
* feat(access-policy): --access-policy enforce, and prove it on a real daemon
`--access-policy` already reached the runtime; what was missing was a way to
ask for deny-by-default without hand-writing JSON, and any evidence that it
works. The README actively said the opposite ("enforcement is not yet
implemented ... a no-op for now") — it has been enforced for a while.
resolveAccessPolicyArg moves out of main.cpp into daemon/access_policy_arg.
so it can be unit-tested, and gains one spelling: the literal `enforce`
expands to {"version":1,"mode":"enforce","restrictions":{}}. That is not a
second switch — `mode` is still the runtime's only switch — it is the bare
document that arms it. Checked before the file branch, so arming enforcement
can't depend on the daemon's working directory.
The integration tests are the point: same binaries, same modules, same call,
policy the only variable. test_ipc_module declares test_basic_module and
test_extlib_module; test_basic_module declares nothing.
no flag -> requestModule(test_basic_module, test_extlib_module) mints
enforce -> the same call is refused, and both names appear in the log
enforce -> requestModule(test_ipc_module, test_basic_module) still mints
The third is the one that matters; a change that refused everything would
pass the second on its own. The refusal is matched structurally rather than
by exact text because the two capability_module implementations in this tree
quote the names differently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(qt-host): link the Qt host runtime from logos-plugin-qt, not logos-qt-sdk
logoscore's daemon and its in-process core service are built on LogosAPI,
LogosAPIProvider and LogosProviderObject. Those moved out of logos-qt-sdk
into logos-plugin-qt, which publishes them as the `logos-qt-host` package
with the CMake target logos-qt-host::logos_qt_host. Point at that target.
Those three headers were the ONLY thing this repo took from logos-qt-sdk —
it emits no Qt consumer wrappers, ships no UI plugin, and never touches
logos_qt_lp_bridge.h or logos_ui_plugin_context.h — so the logos-qt-sdk
input is dropped outright rather than kept alongside. LOGOS_QT_SDK_ROOT
becomes LOGOS_QT_HOST_ROOT in all three derivations (build, tests,
buildPortable), and `--version` now reports the logos-plugin-qt commit.
Both new failure modes are hard errors, never silent skips: an unset
LOGOS_QT_HOST_ROOT is a FATAL_ERROR before find_package runs, and a
find_package that somehow does not define the imported target is a
FATAL_ERROR too.
logos-qt-host needs TokenManager::forIdentity/isolateIdentity, which
logos-protocol only grew on its per-client-token-store commit, so the
lock moves there. logos-plugin-qt is rev-pinned for now because
nix/qt-host.nix does not exist on its default branch yet.
Verified on aarch64-darwin: `nix build .#checks.aarch64-darwin.tests`
passes 21/21 with the committed lock and no overrides (same 21 as the
pre-change baseline), .#cli and .#cli-bundle-dir build, and the set of
LogosAPI/LogosAPIProvider/LogosProviderObject/qtArgDecode symbols in the
logoscore binary is identical to the pre-change build.
* chore(deps): re-pin the SDK stack onto the pushed b4 revs
Rebased onto master, so the inputs have to name the revs the rest of the b4
stack was actually pushed at rather than each input's default branch:
logos-cpp-sdk a04b2788 b3 codegen tip; a strict descendant of
cpp-sdk master, so forward-only
logos-protocol c8bab12 per-client token store — logos-qt-host
calls TokenManager::forIdentity, which
exists nowhere else
logos-plugin-qt cc24fa1 was 8ccb1fc. The superset branch that
logos-liblogos and logos-module-builder
also pin, so exactly ONE logos-qt-host
is in the closure — this CLI links it
directly AND through liblogos_core
logos-liblogos f2a15ef the liblogos built on that same qt-host
logos-capability-module 0cb33fb master, pinned explicitly — see below
All five are rev-pinned in the URL rather than left to the lock: every one is
a branch commit, so an unpinned url lets `nix flake update` silently relock
onto a default branch that does not build here.
capability_module deliberately does NOT move to the universal port (07dba1f).
That port declares metadata.json#host_services and fails closed until a host
calls logos_module_grant_host_services — and nothing in this stack calls it
yet (neither logos-liblogos nor logos-plugin-qt contains a single call site).
Built against it, the daemon's capability gate refuses EVERY requestModule
with "not granted the token_registry host service", so no module can call
another; the new access-policy integration test caught exactly that. 0cb33fb
is what logos-liblogos and logos-standalone-app lock too.
Verified on aarch64-darwin with the committed lock and no overrides:
.#checks.aarch64-darwin.tests-logosctl 191 + 25 + 21 tests, all PASSED
.#checks.aarch64-darwin.tests-logoscore 20 + 24 tests, all PASSED
.#checks.aarch64-darwin.tests built (exit 0)
.#packages.aarch64-darwin.{cli,ctl} built (exit 0)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(deps): rev-pin logos-test-modules at the b4 qt-host tip
The daemon-backed integration checks load these plugins into the daemon
this repo builds, so the two share one host runtime in one process image
-- the same constraint that already rev-pins logos-liblogos. a639b934
links the test modules against logos-qt-host rather than logos-qt-sdk and
carries the matching B4 stack pins; the previous lock sat on master
(f8077fab), which predates that repoint.
The URL had to change, not just the lock. The input was an UNPINNED url,
so it resolved to the default branch -- and f8077fab IS master's tip.
`nix flake update logos-test-modules` was therefore a silent no-op that
would leave the ten b4 commits behind while reporting success.
f8077fab is a strict ancestor of a639b934 (verified on a non-shallow
clone), so this is forward-only, not a lineage switch.
Two behaviour changes ride along and were checked against this repo's
assertions rather than assumed safe:
* test_basic_module and test_extlib_module migrate to
interface "universal". Neither declares metadata.json#host_services,
so the fail-closed gate that keeps logos-capability-module pinned off
its universal port does not apply here.
* stringLength now answers in CHARACTERS, not bytes. Every assertion
here is ASCII ("abcdef" -> 6), so the two agree.
The access-policy fixture still has its pair: test_ipc_module declares
[test_basic_module, test_extlib_module] and test_basic_module declares
none, so basic -> extlib stays undeclared.
Checks built by name, all exit 0: tests-logosctl, tests-logoscore,
tests. 281 tests, 0 failures, 0 skips.
* test: use test_ipc_new_api_module as the transitive-dependency fixture
These integration tests pick a module that DECLARES the other two, so one
load-module has to pull all three, and then request a token across that edge.
test_ipc_module was that fixture; it is being retired as a duplicate. Its
successor declares exactly the same dependency pair, so the fixture role
transfers unchanged.
Worth doing in the same breath as the retirement rather than after: these call
GTEST_SKIP() when the module is missing, so deleting the module out from under
them would not have turned anything red — the dependency-resolution and
token-request coverage would simply have stopped running.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(windows): refuse an unknown target instead of silently skipping it
`logos_use_shared_runtime_from_dll` empties the static archive of each named
IMPORTED target so the symbol resolves to liblogos_core.dll's exported copy
instead. It skipped any name that was not a target, which makes a typo or a
moved target silent — and the failure it hides is the duplicate-statics class:
the image keeps its own static copy of the shared runtime alongside the DLL's,
and PE has no interposition to collapse the two.
That hazard was already WRITTEN DOWN at basecamp's call site ("naming the old
target here would be a silent no-op … Windows would regress to the 29
'rejecting unauthorized call' lines this shim exists to prevent") — documented,
but not enforced. This enforces it.
Taken from feat/sdk-codegen-phase-a, which hardened its logoscore-cli copy and
never fixed basecamp's; feat/sdk-codegen-b3 has neither. It is the one place
where reconciling onto b3 would otherwise lose work, so both copies get it.
Behaviour is unchanged for every current caller: the function early-returns off
Windows, and both call sites pass the same two targets
(logos-qt-host::logos_qt_host, logos-protocol::logos_protocol) that phase-a's
hardened copy already accepts. x86_64-windows still evaluates (386 packages).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: use logos-co/setup-nix-cache-action for Nix setup and caching
Replaces the per-repo installer + cachix pair with the shared action, which
installs Nix with the Logos Attic cache (cache.nix.logos.co) preconfigured and
publishes what the job builds — master to the public cache, every other ref to
ci.
Each converted job also gains
environment: ${{ github.ref == 'refs/heads/master' && 'public-cache' || '' }}
because ATTIC_TOKEN_PUBLIC only exists inside that environment. Without it the
secret resolves empty on master and publishing is silently skipped — the job
still passes, so the omission would not show up as a failure.
The action installs Nix itself on every runner, macOS included. That is a
deliberate reversal of the workaround these files carried: the comments here
said cachix/install-nix-action collides with the runner's pre-existing _nixbld
users (eDSRecordAlreadyExists), so DeterminateSystems' installer was used
instead. It no longer reproduces — logos-delivery-module has already been
converted the plain way and its `build-and-test (macos-latest)` leg passes.
Keeping the workaround would have meant a second installer plus a duplicated
substituter/key block in ten files, guarding against something two green runs
say does not happen. If it ever recurs it fails loudly at install, which is
recoverable; the silent-skip above is the failure mode worth engineering
against.
One property is deliberately NOT carried over: the old cachix step ran with
`continue-on-error: true` so a failed cache push could not fail a job whose
tests passed. The action exposes no equivalent, and adding one here would also
swallow genuine setup failures now that the same step installs Nix rather than
only publishing at the end.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: drop references to removed generator flags and interfaces
README and docs described module authoring in terms of LogosProviderBase,
LOGOS_METHOD and --provider-header, none of which exist. Updated to the
universal model, keeping the retired shapes named as history.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(deps): track master for protocol, cpp-sdk and plugin-qt
logos-protocol#59, logos-cpp-sdk#138 and logos-plugin-qt#19 merged, so the three
rev pins bridging to them are retired, each with its rationale rewritten to name
the PR that closed the gap.
Left pinned: logos-liblogos, logos-capability-module and logos-test-modules —
their branches are still in flight and no merged upstream was confirmed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
e48fc7fef3 |
Add logosctl alongside logoscore: one CLI for daemon, modules and packages (#76)
* feat(core_service): add refreshModules and cascade unload by default Two runtime prerequisites for the logosctl merge. refreshModules() wraps logos_core_refresh_modules(), which liblogos documents as "call after installing new modules so they become discoverable". Basecamp calls it on the package_manager install event, which is why installing a module there needs no restart. core_service did not expose it, so a CLI that installs a package had no way to make the daemon see it short of a restart. unloadModule() now takes withDependents and the CLI defaults it to true (--no-dependents opts out). logos_core_unload_module already accepted the flag; core_service hardcoded false, which left dependents running against an unloaded provider. The result now carries dependents_unloaded so the cascade is reported rather than silent. The dispatch entry defaults a missing second argument to true, so a one-argument unloadModule call keeps working. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(flake): bundle package_manager and package_downloader logoscore bundled only capability_module, so it could authenticate but not manage packages — that was lgpd's and lgpm's job, as separate binaries. Bundling the two package modules is what lets one binary do the whole job. Same trio logos-basecamp bundles, assembled the same way (map the install bundler over the module libs), so the CLI and the GUI drive an identical module surface rather than the CLI being a reduced sibling. Only the package manager ships a distinct lib-portable; the other two are variant-agnostic, matching basecamp's split. Verified against a real daemon: all three are discovered with no module configuration, both package modules load, and package_downloader resolves the live default catalog. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(daemon): make the config dir a self-contained session The daemon knew about ~/.logoscore only as a place to keep its own state files; packages, trust material and persistence lived elsewhere or nowhere. Now the config dir is the whole world for a session: <configDir>/modules installed core modules (writable) <configDir>/plugins installed UI plugins <configDir>/keyring trusted signing keys <configDir>/cache downloaded .lgx <configDir>/data module persistence so copying the directory carries the session's packages and its trust assumptions with it, and two sessions can disagree about both. <configDir>/modules joins the search path beside the bundled dir. Without it an installed module would sit on disk that the daemon could never see, and install-then-load could not work at all. The bundled package modules are loaded at boot and pointed at these directories -- the same four setters basecamp calls -- because every package command is an RPC into them. All best-effort: a daemon that cannot manage packages is still fully usable for loading and calling modules, so none of it aborts startup. Verified live: a bare daemon creates the tree, loads all three modules, and reports the embedded packages via getInstalledPackages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(package): daemon-side install, upgrade and remove Adds the mutating package operations the CLI never had, orchestrated inside the daemon and exposed as core_service.planPackageOperation / applyPackageOperation, plus the `package`, `catalog` and `key` command groups on top. Why daemon-side: package_manager gates destructive work behind a listener-ack protocol with a 3-second deadline. Driving that from a short-lived client would mean holding an event subscription open, interleaving it with outbound calls, and winning a three-second race across the RPC boundary. In-process the ack cannot lose that race, and every client command stays thin and stateless. plan/apply is split so `--dry-run` and the confirmation prompt see exactly what apply will do -- the same dependency-change table basecamp shows, including which running modules get stopped. Without -y and without a TTY the operation is refused rather than assumed-yes, so a script that forgot --yes fails loudly instead of silently uninstalling. install/upgrade take dependencies, remove takes dependents, both by default. Installing never loads: it puts files on disk, and only modules already running beforehand are restarted afterwards. Verified against the live catalog on a portable build: install openmetrics; install chat_module pulling delivery_module in order; re-install as a no-op; install then load with no daemon restart (refreshModules); and removing delivery_module cascading through chat_module with both stopped first. One trap worth naming: LogosList{vec} does not wrap a std::vector the way it wraps a scalar -- it yields an empty args array, and the module sees a zero-argument call it cannot dispatch. The batch uninstall builds its argument explicitly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(config): replace the flag surface with a YAML session document Configuration was ~20 flags plus two hand-written mini-grammars: a `NAME=PROTOCOL[,k=v...]` parser that existed only to squeeze a nested structure through a flag, and a per-flag defaults<config<CLI merge. Both are gone. main.cpp drops from 943 to 531 lines. Configuration now lives in the session: <configDir>/daemon/config.yaml written by `daemon config set` <configDir>/client/config.yaml written by `client config set` and is never passed alongside an unrelated command, so `daemon start` and every client command take the session exactly as it is on disk. --config-dir is the one surviving flag, because it selects *which* session to act on and so cannot itself live inside one. The split is by audience: files a human edits are YAML, files the daemon and modules own stay JSON (state.json, tokens, the auto token). Converting through nlohmann::json means the existing validated daemonConfigFromJson / clientStateFromJson keep doing the schema work. Two traps fixed while wiring it up, both of the accept-then-ignore kind that leaves an operator with no explanation: - A bare `modules: {core_service: [ ... ]}` sequence was silently skipped (only the `{transports: [...]}` spelling parsed). It is now accepted as shorthand. - Unknown top-level keys are rejected by `config set` and the error names the correct spelling, so `insecureTcp` no longer looks like it worked when the key is `insecure_tcp`. Module search paths remain configurable via the `modules_dirs` key, which is what replaces -m for tests and dev loops. The eight CLI tests that covered deleted flags are rewritten against the new surface: malformed YAML rejected without clobbering the existing config, unknown keys named, set/show round-trip, and absent config treated as defaults rather than an error. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(cli): logosctl, with docker-style command groups Renames the binary and reorganises ~20 flat hyphenated commands into groups: daemon, client, module, token, package, catalog, key, plus top-level aliases for the four verbs that cannot be confused with a runtime module (status, call, watch, stats) and the two package verbs with no module meaning (install, search). The hyphenated names survive as internal dispatch tokens but are hidden from --help: `module load X` is rewritten to `load-module X` in argv before CLI11 parses. The rewrite happens in argv rather than via nested CLI11 subcommands because daemonSub->fallthrough() pushes a nested subcommand's unmatched arguments up to the top level, where they are rejected ("The following argument was not expected: show"). `module` is no longer an alias for the verbose call syntax -- it is the group. Use `call`. Also implements --detach, which was specified but missing. It re-execs rather than continuing in the forked child: macOS refuses to let a process that has already initialised CoreFoundation keep running after fork(), and the Qt/liblogos link pulls CoreFoundation in before main. The child redirects stdio to <configDir>/daemon/daemon.log -- without that the shell never sees EOF and `daemon start --detach` appears to hang -- and the parent returns only once state.json exists, so the next command cannot race the boot. Env vars and the default session directory rename to LOGOSCTL_* and ~/.logosctl. User-facing messages now name the group grammar rather than the internal tokens. Verified on the portable build: daemon start --detach returns in ~3s with a working daemon; catalog ls, search, install --dry-run, install, package ls, module ls/load/show, upgrade (no-op), and remove of a loaded module all behave. 18/18 CLI tests, unit tests unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: rewrite for logosctl, sessions and YAML config The README documented a flag surface that no longer exists (--persist-config, --module-transport, --modules-dir, the seven --client-* flags) and had no account of sessions or packages at all. Replaces the daemon/transport/persist-config sections with: what a session directory is and why it is portable, `daemon config set` and the YAML schema, and the package/catalog/keyring commands. Keeps the two hard-won warnings that are still true -- a remote daemon must expose capability_module as well as core_service, and plaintext tcp on a non-loopback host needs an explicit opt-in. Doctests are renamed and rewritten around sessions: the daemon spec no longer passes -m but seeds ./session/modules, and uses `daemon start --detach` instead of backgrounding with & (which returned before the transports bound and raced the first command). Also fixes the stats table: the MODULE column was a fixed 12 characters, so a real name like "test_basic_module" ran straight into the PID with no separator. Verified the rewritten daemon-doctest sequence by hand against a dev build: seed session, start --detach, module ls/load, call, stats, stop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * doctests: use --detach and the session log `daemon start --detach` already returns only once the daemon is accepting commands and sends its output to <session>/daemon/daemon.log, so the `sh -c '... > logs.txt 2>&1 &'` wrapper is not just redundant -- it hid the output the specs then tried to cat. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: ship logosctl alongside logoscore instead of replacing it logosctl is new and unvalidated; logoscore is what people depend on today. Replacing one with the other in a single step meant every consumer had to move at once, on trust. Shipping both means logosctl can be validated in real use first, and logoscore removed afterwards. Both binaries build from this repo over one shared runtime -- daemon, core_service, client, output. They differ only in main.cpp and a Config::Flavor the front-end sets, which selects the config directory, the env var consulted for an override, the config file names, and the format they are written in. The isolation is the point, so it is deliberate and tested: logoscore ~/.logoscore LOGOSCORE_CONFIG_DIR daemon/config.json logosctl ~/.logosctl LOGOSCTL_CONFIG_DIR daemon/config.yaml A logosctl session cannot disturb a logoscore deployment. Reading needs no branch -- YAML is a superset of JSON, so one parser handles both -- only writing differs. logoscore is behaviourally unchanged, which took two specific decisions: - The session directory and the package-module bootstrap are gated on the modern flavor. Auto-loading two extra modules would change what `status` and `list-modules` report, and logoscore's doc-tests assert those exact counts. - The bundled package modules live in modules-pkg/ rather than modules/, because logoscore scans the latter and would otherwise report two modules it never had. Verified: `logoscore --help` is the old flat surface with all four flag families intact; a logoscore daemon reports loaded:1 not_loaded:0 and creates no session directories; both daemons run at once with separate state. Its doc-tests are restored unchanged. logosctl gets its own, including a new logosctl-packages spec covering the capability that motivated the merge -- search, dry-run, install, load, remove -- verified end to end against the live catalog. 122/129 unit tests, 18/18 CLI tests. The 7 OutputTest failures are pre-existing on master. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build: give logosctl its own flake outputs Building both binaries into one package meant `nix build` and `.#cli` started handing out logosctl too, which is the opposite of keeping the two apart while the new one is validated. Now each output ships exactly one binary: .#cli .#cli-bundle-dir .#cli-appimage -> logoscore .#ctl .#ctl-bundle-dir .#ctl-appimage -> logosctl .# (default) -> logoscore So anything already pointing at the default or at `.#cli` -- including every doc-test across the workspace that does `nix build github:logos-co/logos-logoscore-cli` -- keeps getting the tool it gets today, and logosctl is strictly opt-in. They still compile together, since they share everything but main.cpp; only the packaging is split. modules-pkg/ ships solely in the ctl outputs, because logoscore never scans it. logoscore's desktop entry and icon are restored, and logosctl gets its own. The logosctl doc-tests now build .#ctl / .#ctl-bundle-dir. Verified: every output builds and ships only its own binary; both portable bundles run and report the module set expected of each. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(config): let each session subdirectory be redirected The session directory being self-contained is what makes it portable, and that should stay the default -- but it was also the only option, which made reasonable setups impossible: sharing one keyring across sessions, putting the .lgx cache on a bigger disk, or pointing at a modules tree something else manages. A `dirs:` block now redirects any of them: dirs: keyring: ~/.config/logos/trusted-keys cache: /var/cache/logos modules: /opt/logos/modules plugins: plugins-custom data: /var/lib/logos/data The form of the value decides whether portability survives, which is the part worth knowing: plugins-custom -> <session>/plugins-custom still portable ~/x -> $HOME/x outside the session /var/cache/... -> as given outside the session `~` is handled because it is the natural thing to write in a config file and would otherwise resolve to <session>/~/... , which exists nowhere. Overrides resolve once, when set, so relocating a session afterwards cannot silently drag an absolute path along with it. Only the daemon applies them, and it does so before anything asks Config for a path. persistence_path is folded into dirs.data -- it was already the same setting under an older name -- so the two no longer need choosing between at the point of use. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(daemon): rotating log file with configurable size and retention --detach used to dup2 stdout/stderr straight onto daemon/daemon.log, which grew without bound and had no rotation. A long-lived daemon needs better than that. There is now a logs/ directory, like Basecamp's, and a logging block: logging: enabled: true # false -> no log file at all file: daemon.log # inside dirs.logs max_size_mb: 10 # rotate past this; 0 = never rotate max_files: 5 # keep this many in total console: true # mirror to the terminal dirs.logs joins the overridable session directories, so logs can be shipped somewhere a collector already watches. Capture is pipe-based rather than a file redirect, and that is the load-bearing decision: module hosts are separate processes holding inherited descriptors. Redirecting to a file catches their output but makes rotation impossible -- renaming a file out from under a child that has it open just keeps filling the old inode. A pipe puts one reader in charge, so rotation is safe and subprocess output still lands in the log. Same shape as basecamp's LogRedirector, which solved this already. The size cap and retention come from spdlog's rotating sink rather than being hand-rolled; liblogos already logs through spdlog. Lines arriving from the pipe already carry their own timestamp and level, so the sink uses a raw pattern instead of stamping them twice. Two bugs found while testing it: - Draining raced shutdown. stop() cleared the running flag before restoring the descriptors, so a reader holding data would process it, loop, see the flag clear and exit -- dropping whatever was still in the pipe. The last lines before a shutdown are exactly the ones worth keeping. EOF is now the only stop condition. - --detach reported the wrong path. The parent prints before the child has read the config, so it guessed the default and lied to anyone who had redirected dirs.logs or renamed the file. It now reads the same config the child will. Verified live: default, disabled, and redirected-with-custom-filename all behave and are reported accurately. Four unit tests cover capture of both streams, no double-stamping, rotation with retention, and disabled-is-not-an-error. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(daemon): timestamp log files, and bound the directory Adopts basecamp's naming -- each start writes its own daemon_<yyyymmdd_HHMMSS>.log -- so a session's output is one file you can point at, instead of every run appending into the same daemon.log. Two things beyond copying basecamp: - `logging.file` survives as a symlink to whichever file is current, so `tail -F logs/daemon.log` follows across restarts and nobody has to work out a stamp. It also means --detach can report a path that is always valid; previously it had to guess one, and guessed wrong for anyone who had redirected dirs.logs. - max_files now bounds the *directory*, pruning oldest-first at each start. spdlog's retention only prunes within one sink's rotation set, and every start opens a new stamped base name, so without this a daemon restarted a hundred times would leave a hundred logs behind. basecamp has exactly that problem. Verified live: three restarts leave three stamped files with the symlink tracking the newest; five restarts with max_files: 2 leave two. Two new tests cover the naming and the symlink resolving to the current session, and the cross-session pruning. The rotation test needed fixing too -- it counted the symlink as a log file, which predated the symlink existing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: ignore suffixed nix out-links .gitignore listed `result` but not `result-*`, so every out-link from a targeted build -- `nix build '.#ctl' -o result-ctl`, `-o result-tests`, and so on -- was untracked-but-not-ignored, and `git add -A` committed them as symlinks into /nix/store. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: keep releasing logoscore, and release logosctl beside it The earlier rename left the release workflow building the `cli-*` outputs -- which are logoscore -- while naming every artifact `logosctl-*`. A release/** push would have shipped logoscore binaries under the wrong name, and stopped releasing logoscore under its own. Both are now built and published as separate, correctly-named assets: logoscore-{x86_64,aarch64}-linux.tar.gz from .#cli-appimage logoscore-aarch64-macos.tar.gz from .#cli-bundle-dir logosctl-{x86_64,aarch64}-linux.tar.gz from .#ctl-appimage logosctl-aarch64-macos.tar.gz from .#ctl-bundle-dir logoscore's asset names are exactly what they were, which matters: release sets fetch this repo and expect a bundle containing `bin/logoscore`. Each tool builds from its own flake outputs, so an asset labelled logoscore contains logoscore and nothing else. Both jobs gained a tool matrix with fail-fast disabled, so a failure in the under-validation logosctl cannot block a logoscore release. The release job now collects artifacts by pattern instead of naming each one, so retiring logoscore later means deleting a matrix entry rather than unpicking a download list. Release notes lead with logoscore as the tool to use, and say the two share no state so installing logosctl cannot disturb an existing setup. The doc-tests workflow globs doctests/*.test.yaml, which now covers both suites, so it is no longer named after one of them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: run both tools' suites, in parallel While both binaries ship, both get tested. logoscore had no automated coverage on this branch at all -- only its doc-tests -- so a change to the shared runtime could regress the tool people actually use and nothing would say so. tests/test_cli_logoscore.cpp and tests/test_integration_logoscore.cpp are copies of the suites frozen against logoscore's surface. Copies rather than a parameterised shared suite on purpose: the two surfaces genuinely differ, and this way retiring logoscore is a delete rather than an unpick. checks.tests-logosctl and checks.tests-logoscore are separate derivations, so nix builds them concurrently; checks.tests aggregates both, keeping `nix build .#checks.<sys>.tests` working for CI while now covering both tools. It immediately earned its keep, catching three regressions: - The integration harness still passed -m, which logosctl no longer accepts, so its daemon never started and seven integration tests were failing on this branch. It now writes the modules_dirs config the daemon reads. - `logoscore --version` reported "logosctl version ...". The version banner had been renamed wholesale; each front-end now names itself. Exactly the sort of thing nobody notices until a bug report cites the wrong tool. - The new log sink only mirrored to the console when stdout was a TTY, so `logoscore -D > logs.txt` -- which the doc-tests do -- produced an empty file. Mirroring now follows the configured setting, pipe or terminal alike, and the log file is gated to logosctl so logoscore's output behaviour is untouched. Both suites green: logosctl 138 unit + 18 CLI + 18 integration, logoscore 20 CLI + 18 integration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(package): honour -o, and stop parsing command lines backwards `package download -o DIR` accepted the flag and threw it away -- the argument was parsed into a variable and then explicitly discarded with `(void)outDir;`. The file went to $TMPDIR regardless. The config's `dirs.cache` had the same problem from the other end: the directory was created and documented as holding downloads, and nothing ever wrote to it. The cause was the same for both. package_downloader takes no destination, so the file lands in $TMPDIR on the DAEMON's filesystem -- which is where the move has to happen too. Doing it client-side would work only for a local daemon. So `downloadPackage` joins the daemon-side package operations: it downloads, then moves the result into the requested directory, or into the session's cache/downloads when no -o was given. The client resolves a relative -o against its own working directory first, so a local daemon does what the user typed; against a remote one the path is remote, and a bad one fails loudly rather than quietly writing elsewhere. Writing the first test for it turned up something worse. CLI11's `parse(std::vector<std::string>&)` consumes the vector from the BACK -- only the rvalue overload reverses for you -- so passing natural order parses the command line backwards. `watch` and `issue-token` did reverse first; nothing else did. It goes unnoticed with one positional and flags (order does not matter), and is quietly wrong the moment an option takes a value, because the option pairs with the token to its LEFT: package download pkg -o dir -> name="dir", output="pkg" package install a b --version 1.0 -> names=["1.0","b"], version="a" So `package install`, `search --category`, and `download -o` all misparsed. Every site now goes through one `parseArgs` helper that reverses, which fixes the broken ones, is a no-op for the harmless ones, and removes the trap for the next command. PackageCommand had no unit tests at all, which is why a discarded flag survived review. Four now cover download; the two asserting -o reaches the daemon fail against the old code. 142 unit + 18 CLI + 18 integration green for logosctl, 20 + 18 for logoscore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: one README about the repo, one document per tool The README had grown into a logosctl manual with a banner on top telling logoscore users that everything below did not apply to them, and pointing them at doc-test YAML for their actual documentation. Since logoscore is still the tool to use, its documentation should not be the thing you are told to skip. So: README.md covers what is true of both -- what the repo is, the two binaries and how they differ, the flake outputs, the test targets, dependency resolution, platforms -- and hands off to one document per tool. docs/logoscore.md the usage material, unchanged, as its own document docs/logosctl.md sessions, config, logs, packages, examples Writing logosctl's own document exposed a gap: it had no command reference at all. The rewrite dropped the client-command list, argument typing and exit codes, and left behind a "see Argument typing below" pointing at a section that no longer existed. All three are back, with the command list written against the grammar that is actually implemented (checked against normalizeGroupVerbs and the subcommand dispatch, not from memory), plus the two defaults worth stating up front -- install does not load, remove takes dependents. Also fixes stale copy that survived the earlier rewrite: `load-module` where logosctl says `module load`, and a "multiple module directories" caption over a --config-dir example, from a flag logosctl does not have. Deleting logoscore later is now deleting one file and a table row. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(daemon): make TLS configurable again, and say why startup failed Three bugs, all found by running the doc-tests I had just rewritten instead of trusting them. **tcp_ssl could not be configured at all.** `transportFromJson` never read `cert` or `key`. That was harmless while those arrived via `--module-transport ...,cert=...,key=...`, parsed by the CLI mini-grammar -- but that grammar is gone, and the config file is now the only place to set them. So every tcp_ssl listener bound with no certificate: the daemon started, reported itself healthy, accepted connections, and failed every handshake with "no shared cipher (SSL routines)". The client just saw "core_service not reachable". The stripping was deliberate but applied one layer too high: cert and key have no business in state.json, which clients read, but the config file is where an operator *authors* them. `transportToJson` now takes `includeSecrets` -- true writing the config, false writing state.json. A test asserts the round-trip, and another asserts the key path never appears in state.json. **`--detach` swallowed the reason startup failed.** Config validation runs before LogSink opens the log, and the child's stderr went to /dev/null, so a rejected config produced "daemon exited during startup. See <path>/logs/daemon.log" -- naming a file that had never been created. The child's early output now goes to a startup file the parent reads and prints on failure, removed either way. LogSink takes those descriptors over as soon as it starts, so the file only ever holds pre-logging output. **The plaintext-TCP guard advertised a flag that does not exist.** It said "pass --insecure-tcp"; logosctl has no such flag. It now names the config key, `insecure_tcp: true`. Verified end to end against a real daemon: plaintext guard refuses and says why, loopback TCP binds and serves `status`/`module ls` from a separate client session, TLS serves the same over 6443/6444, and dropping the CA while keeping verify_peer still fails closed. 144 unit + 18 CLI + 18 integration green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(flake): give autoPatchelf the libraries both binaries now link Every Linux build failed: auto-patchelf could not satisfy dependency libyaml-cpp.so.0.8 wanted by .../bin/.logoscore-wrapped The packaging derivations listed only Qt in buildInputs, which is what autoPatchelfHook resolves DT_NEEDED entries against. yaml_json.cpp and the log sink are in the shared sources, so *both* binaries link yaml-cpp and spdlog -- including logoscore, which is why its Linux build broke too on a branch that was supposed to leave it alone. macOS does not patchelf, so this was invisible locally and in the macOS CI jobs; only the Linux matrix caught it, and it took down the AppImage builds, the CI job, and every Linux doc-test with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(doctests): bring the logosctl specs up to what logosctl does Nineteen doc-test steps were failing. None of them were runtime bugs in the specs' own right -- they were specs still describing an older logosctl, which is its own kind of failure: a doc-test that lies is worse than no doc-test. transports Still drove `--module-transport` and hand-written client/config.json. The flags had been dropped from the `run:` lines but no config step replaced them, so the daemon never bound TCP at all and every step after it failed. Rewritten around `daemon config set` / `client config set` with YAML documents, for both the plaintext and TLS halves. daemon Read the log at session/daemon/daemon.log; logs moved to session/logs/. The crash-recovery step passed `-m`, which logosctl does not accept, so its daemon never started and the step reported LEAKED against a worker that had never existed. modules-bundle Asserted all three modules in result/modules. The package modules live in modules-pkg/ so that logoscore's modules/ stays byte-identical -- which the spec is now the place that explains. packages Expected the interactive wording ("dry run", "Installed:"). Doc-tests are not a terminal, so every command renders JSON. The install was working the whole time; only the assertions were wrong. They now match the JSON, and the prose says why it is JSON. Rewriting the transports spec is what turned up the TLS and --detach bugs fixed in the previous commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): a typo must not abort the daemon, and a key must not lie Three defects in the YAML config path, all found by building a Python client against this CLI and checking its assumptions against the binary rather than the docs. **A config typo aborted the process.** printf 'version: 2\nmodules_dirs: /single/path\n' > bad.yaml logosctl --config-dir ./s daemon config set ./bad.yaml => libc++abi: terminating due to uncaught exception ... [json.exception.type_error.302] type must be array, but is string nlohmann's `json::value(key, default)` THROWS when the key is present with the wrong type, every config read used it, and nothing caught it. So it was not one key -- it was every key in both readers. A scalar where a list belongs is an ordinary mistake and it killed the binary. Now a type-checked reader (src/json_schema.h) records "<dotted.path>: expected <what>, but got <what>" and the document is refused whole, the same shape as the existing unknown-key error: {"code":"INVALID_CONFIG", "message":"modules_dirs: expected a list of strings, but got a string."} Both readers went through it, including two paths that could abort the daemon mid-boot rather than at `config set`. **`config set` validated after writing.** A schema-invalid document was installed and then reported as an error, leaving the session holding a config the daemon would refuse to boot from. Validation now happens entirely in memory first, on both the daemon and client sides -- the client side had no schema validation at all -- and the write is temp-file + rename instead of truncate-in-place. That exposed a fourth: `yaml_json::dump` emitted numeric-looking strings bare, so `port: "6001"` came back as the number 6001. The bytes validated were not the bytes written. **Two keys were accepted, stored, and never applied.** `signature_policy` sat on the allowlist and was written verbatim to config.yaml but was never even parsed. An operator setting `require` got no enforcement and no warning. It is now parsed with a strict allowlist and pushed into package_manager at boot beside setKeyringDirectory -- the module has had setSignaturePolicy all along. Unset issues no RPC, so the module keeps its own default instead of having it restated. The top-level `ssl: {cert, key, ca}` block was parsed into DaemonConfig and read by nobody; only per-listener cert/key reached the transport set. Configuring TLS the obvious way therefore produced listeners with no certificate and "no shared cipher" on every handshake -- the same failure fixed one layer down last commit. It is now a session-wide default that per-listener values override. Also: docs advertised `module load --no-deps`, which does not exist -- `module load` takes only a positional name and always resolves dependencies. Corrected, along with the rest of the command reference, verified against the binary. logoscore is untouched: 20 CLI + 18 integration, exactly as before. logosctl 171 unit (was 144) + 25 CLI (was 18) + 18 integration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(detach): re-exec the launcher, not the ELF it hides `daemon start --detach` was dead on Linux portable builds. The daemon exited immediately with status 127, no output, and no log file -- so the only diagnostic was "daemon exited during startup. See <path>", naming a file that had never been created. strace, on a real Linux box, said it in one line: execve(".../bin/.logosctl.elf", [...]) = -1 ENOENT exit_group(127) A portable bundle installs the CLI as a launcher script beside a hidden companion: bin/logosctl the launcher, a shell script bin/.logosctl.elf the real ELF The launcher exists because that ELF cannot be started on its own: its PT_INTERP names a dynamic loader that is not on the host, so the launcher runs it through a known-good ld.so instead. The ENOENT is the kernel reporting the missing *interpreter* -- the ELF is right there. --detach re-execs itself (it has to: macOS forbids running a forked process that has initialized CoreFoundation), and it re-exec'd executablePath(), which is that ELF. My first attempt preferred argv[0], reasoning that it is what the caller actually typed. That was wrong, and the trace showed it failing identically: the launcher execs ld.so with the ELF, ld.so drops itself from argv, and the program sees the ELF as argv[0] too. Neither source of truth names the launcher. So the mapping is applied to whatever candidate we end up with, using the convention the launcher script itself documents -- the install dir is the one holding the hidden companion `.$BASE.elf`. `bin/.logosctl.elf` maps back to `bin/logosctl`. argv[0] is still preferred over executablePath() (it is what was invoked, and it is right when a bare name resolves through PATH), and it is absolutised, since the daemon may run from a different directory. Only this combination was ever broken: portable AND Linux AND --detach. macOS bundles a real binary with qt.conf and no launcher, Linux dev builds are ordinary ELFs, and the foreground -D path never re-execs. The one doc-test that uses the portable bundle is the packages spec, and cachix served a permanent 522 for one of its store paths from the day it was written -- so its 14 cascading failures read as infrastructure until the cache recovered and the real failure surfaced underneath. Verified on Linux against the same bundle that failed: daemon starts detached, all three bundled modules load, `daemon stop` returns ok. Also here, and what made the diagnosis possible: --detach now prints the TAIL of the daemon log rather than its path. The startup file only holds output from before LogSink takes the descriptors, so a daemon that dies after logging is up left it empty and the reason unread. That there was no log at all is what pointed at exec. 179 unit tests (8 new, covering the launcher mapping and its edges: no sibling, an ordinary foo.elf, a non-executable candidate, absent argv[0]). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * build: bundle logosctl/logoscore as headless Qt programs * build: bump nix-bundle-dir and nix-bundle-appimage to main Picks up the merged trampoline drop: per-arch psABI PT_INTERP, DT_RPATH, qtCliApp for headless Qt, and the AppImage consumer that already tracks the same pin. nix-bundle-dir 4fd87d1 (PR tip) → cb9afc8; appimage 8fcc56b → 04a3cf8. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
5623313431 |
fix: reap stale socket files at daemon boot (#71)
The reported symptom was dozens of lingering /tmp/logos_* files. Two
separate paths produce them, and only one was fixed:
* A graceful stop now unlinks its own sockets — the SIGTERM self-pipe
handler lets QCoreApplication::exec() return, so QLocalServer's
destructor (the only thing that unlinks) actually runs. That landed
in logos-module-loader-qt#4 and is what eliminated the bulk of the
volume, since previously *every* shutdown leaked one file per module.
* Nothing cleans up after a hard kill. SIGKILL, a module crash, and the
PR_SET_PDEATHSIG kill that reaps orphaned logos_host children all
skip destructors by construction, so those files accumulate forever.
logos::reapStaleSockets has been available (and gtested) in logos-protocol
since #20 but had no production caller. Wire it into daemon boot.
Placed after the already-running check and before logos_core_init: at
that point none of our sockets exist yet, so the reaper cannot race its
own endpoints. Co-resident nodes are safe because logos::isSocketDead
fails closed — it unlinks only an S_ISSOCK inode that we own and whose
connect() is refused — so a live socket and a regular file sharing the
prefix are both left alone. QDir::tempPath() is the authoritative
directory: QLocalServer resolves bare server names against it, and it
honours $TMPDIR on Linux and macOS alike.
Two integration tests, in a private $TMPDIR so the assertions describe
this node rather than a dev box's leftovers:
NoSocketsSurviveGracefulStop — loads a module (so the tally includes a
logos_host child socket), asserts sockets exist, stops, asserts none
remain. The "sockets exist" precondition is asserted so the test cannot
pass vacuously.
BootReapsStaleSocketsButSparesLiveOnesAndFiles — plants a stale socket
(bound, listener closed), a live socket held open by the test process,
and a regular file named logos_execution_zone-1.0.0.lgx, then boots a
daemon and asserts only the stale one is gone. The last two are the
fail-closed guarantees: reaping the live socket would break a
co-resident node, and a glob-based cleanup would delete a
multi-hundred-MB build artefact that really does sit in the temp dir.
Verified by negative control: with the reapStaleSockets call disabled,
BootReapsStaleSockets... FAILS and the rest stay green. Full suite 18/18
on aarch64-darwin.
Socket paths are capped by sockaddr_un::sun_path (104 bytes on macOS),
and macOS's default temp dir is already ~50 chars, so the fixture picks a
directory short enough to hold "/logos_capability_module_<12 hex>" and
skips if none fits — otherwise every listen() fails with a confusing
HostNotFoundError.
|
||
|
|
679a9af8fd |
fix: report module version in list-modules and module-info (#59) (#60)
* fix: report module version in list-modules and module-info (#59) The version column was always empty and the JSON omitted version entirely (`delivery_modulev` in the table = name + "v" + empty). The data layer never populated it. Source it generically from liblogos' new logos_core_get_modules_info(), which returns name/path/loaded/dependencies/dependents/metadata per known module. listModules and getModuleInfo now build from that single call, so: - list-modules shows VERSION (table) and "version" (JSON) for loaded AND not_loaded modules (version comes from on-disk metadata). - module-info reports version plus dependencies/dependents. - load-module / reload-module responses include the version. Output: empty versions render as "-" (modules aren't required to declare one), and the NAME/VERSION table columns size to their content so long names no longer collide with the version. Tests: OutputTest cases for the table/dash/collision rendering; an integration ReportsModuleVersion test (real daemon) covering version + dependencies across list-modules/module-info/load-module; doctest assertions. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat: report uptime for loaded modules (#59) list-modules, status, and module-info now report uptime_seconds for loaded modules, derived from the load timestamp liblogos records (now - loaded_at). Unloaded modules report no uptime_seconds (uptime is loaded-only). The daemon stamps loaded_at with the same wall clock core_service reads, so the value is consistent. Tests: ReportsModuleVersion asserts uptime_seconds is absent for an unloaded module and present once loaded. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: bump logos-liblogos to merged modules-info API (#159) Re-pin logos-liblogos 87ae7ce → 819faac (master, #159) and its transitive logos-module a3e288a → 2ec64c4 (#21), so the version/uptime/deps features build against the merged generic modules-info API without overrides. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat: add --no-json/--human flag; doctest shows both output forms Every client command auto-selects human output on a TTY and JSON when piped. Add --no-json (alias --human) as the explicit inverse of --json, forcing the human-readable form even when piped — useful for scripts/log capture and for docs that want to show the terminal view deterministically. The "Running Modules with the logoscore Daemon" doctest now shows both the human table and the JSON for status, list-modules, module-info, stats, and call, using --human/--json to render each form. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(doctest): show human and JSON output samples for both forms The doctest generator renders commands and prose but not captured output, so add post_text blocks displaying both the human-readable and JSON output for status, list-modules, module-info, stats, and call. The steps already run both forms (--human/--json) and assert on them; this surfaces the outputs in the generated tutorial. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(doctest): restore unrelated generated outputs run.sh clears outputs/ wholesale and only regenerates the spec(s) passed to it, so regenerating just logoscore-daemon.md inadvertently dropped the transports and concurrent-blocking tutorials. Restore them unchanged — this PR only touches the daemon doctest. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bfb64580b0 | fix: add core_service token | ||
|
|
c410a01091 |
replace QJson* with nlohmann/json
replace QJson* with nlohmann/json
update flake
ci fix
update flake
update flake
fix include path to use SDK's include/cpp where logos_api_client.h actually lives
The combined SDK symlinkJoin package places headers at include/cpp/ (from
the headers sub-package), not at include/ root. Using include/cpp ensures
the compiler finds the SDK's logos_api_client.h (with nlohmann::json overloads)
before the stale copy shipped inside logos-liblogos's include directory.
Co-authored-by: Cursor <cursoragent@cursor.com>
fix watch command args order for CLI11 parse
CLI11's parse(vector<string>) processes from the back of the vector, so
args must be reversed before calling parse(). The
|
||
|
|
1ba40efead | ci fix | ||
|
|
ab7b2179bf | replace QString | ||
|
|
823fa438bc |
test module subprocess crash scenario (#29)
* test module subprocess crash scenario * bump liblogos |
||
|
|
454e0696e9 |
add integration tests (#27)
* add integration tests * PR comments |