14 Commits
Author SHA1 Message Date
Dario LipicarandClaude Opus 5 e9f134f5aa fix: draw the line at LOADED, not merely known — for both watch and call (#109)
* fix(watch): defer the subscription instead of refusing a module that is not up yet

`logoscore watch <module>` answered WATCH_FAILED and exited whenever the
module's registry socket had no listener at that instant, and nothing ever
retried. Two ordinary states hit it:

  * subscribing before the module is loaded at all;
  * subscribing in the window after `load-module` RETURNS but before the
    module publishes its object. `load-module` answers once the plugin is
    in; publishing happens afterwards.

Measured on macOS, cold first load of a freshly built plugin:

    .803  Registry connect attempt started (no peer contact yet)
    .804  LogosAPIConsumer: Requesting object: "modules_state"
    .804  Warning: Not connected to registry. Cannot request object
    .957  [modules_state] RemoteTransportHost: Published object

Refused at .804; published 153 ms later. The refusal is
LogosAPIConsumer::requestObject's `isConnected()` guard --
`m_connected && endpointHasListener()`, a socket liveness probe. This is the
third un-migrated instance of the one-shot subscribe shape; the generated Qt
wrappers and the QML bridge both moved to onEventWhenAvailable already.

The dangerous part was never the failure, it was that nothing checked it: a
script's only event assertion went dead while every later line still passed.

It is also inconsistent with the call path, for identical input. Same module,
same instant, not loaded:

    watch  -> WATCH_FAILED in    139 ms
    call   -> RPC_FAILED  in 20 544 ms

invokeRemoteMethod reaches the transport through acquireCachedObject and never
sees the guard, so it waits out waitForSource(20s). One path gave up in a
millisecond, the other waited 20 seconds.

Named events now go through onEventWhenAvailable, which holds the subscription,
arms it when the module appears (including one loaded later in the session) and
re-arms across a reconnect. The wildcard form (`watch <module>` with no
--event) cannot: it subscribes with an EMPTY event name, which LogosObject
reads as "every event" but onEventWhenAvailable refuses outright -- so a fix
covering only the named form would have left the CLI's DEFAULT invocation
silently dead. It goes through whenObjectAvailable instead.

Deferring must not swallow a typo, so a name the host has never heard of still
fails fast rather than parking on a subscription with no future. Before this
change both cases failed identically and that was free; now it is pinned.

Three tests in each integration suite, against a fixture whose module is
discoverable but NOT loaded -- deterministic, where reproducing via a
post-load-module subscribe needs a cold machine. Verified as a negative
control: with the tests present and this fix reverted, both "arms later" tests
fail with the exact production signature,

    {"code":"WATCH_FAILED","message":"Failed to watch events for module
     'test_basic_module'.","status":"error"}

and WatchUnknownModuleStillFailsFast passes, confirming it pins existing
behaviour rather than the new path. With the fix: checks.tests-logoscore and
checks.tests-logosctl both green (250 + 30 + 27 tests), the two arming tests
dropping from 64s/70s of failing poll loop to 1.8s/1.9s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(watch): draw the line at LOADED, not merely known

Follow-up to the previous commit, which gated on getKnownModuleNames() and so
deferred a subscription for any module the host had ever heard of. That made
`watch` on an unloaded module park forever on a subscription that might never
arm — measured here as `timeout 12` reporting 124, with no answer printed at
all. A module that is not loaded may never be, so refusing is the useful
answer, and a typo now gets that same answer rather than a worse one.

The contract is two halves:

  * NOT LOADED -> fail fast. Already true before any of this; now pinned.
  * LOADED     -> always succeed. This is the half that was broken, and the
    reason the deferred path is here at all: `load-module` returns once the
    plugin is in, but the module publishes its object afterwards, so there is a
    window where the host reports a module loaded and the one-shot
    requestObject() refused it. Cold, that window is seconds.

Past the loaded check the subscription cannot fail, by construction:
onEventWhenAvailable / whenObjectAvailable answer 0 only for arguments they
refuse (empty name, null callback), none reachable here, and neither touches
the transport on the calling thread.

COVERAGE, stated plainly because it got weaker and the weakness is not visible
from a green run. The unloaded half is deterministic — the fixture's daemon
starts with the module discoverable but not loaded, and that state holds still.
The loaded half is NOT, and cannot be made so from the CLI: nothing can hold a
module in "loaded but not yet published". Verified rather than assumed — with
these tests against the ORIGINAL master implementation, all three PASS on a warm
machine. They bite on a cold CI worker, which is where the failure was seen,
but they are not a warm-machine regression net.

Two deterministic alternatives were considered and rejected. Subscribing before
the load reproduced it reliably, and was what the previous commit tested, but it
is no longer valid behaviour under this contract. Crashing the module leaves it
briefly listed as loaded with its socket gone — the right state, reached by a
race, since the daemon's SIGCHLD handler flips it out of "loaded"
asynchronously. A racing test is worse than none.

checks.tests-logoscore and checks.tests-logosctl both green (250 + 30 + 27).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(call): draw the line at LOADED, not merely known

The same line the previous commit drew for `watch`. `callModuleMethod` had no
precondition at all: its `!moduleClient` guard can never fire, because
LogosAPI::getClient constructs a client for any name and never returns null.

So an absent module was found the expensive way, and two 20s deadlines raced to
report it — the daemon's acquire (object_unavailable) against the CLI client's
own RPC deadline (RPC_FAILED, a client-side literal the daemon cannot emit).
Measured: 17 calls split 5/12 across the two codes, all at 19.99-20.43s.

Not fixed by a shorter deadline: one Timeout feeds both the acquire and the
call, and the long acquire is the documented startup contract. The daemon
already has an in-process answer and already uses it in watchModuleEvents.

core_service is exempt (published by the daemon's own provider, never in the
loaded set). The startup window is unaffected: liblogos marks loaded before
publish, so a warming module passes the gate and keeps its full budget.

Tests assert the property, not a tally — the old behaviour was bimodal per run,
so sampling can pass unfixed. A positive control covers the load->publish
window, which is what a shortened deadline would break. Two assertions in
NoLoadNegativePaths were passing BY hanging (timeout's 124 is non-zero); both
now exclude it.

Follow-ups, not in this commit: logos-test-modules conformance still expects
["object_unavailable","RPC_FAILED"] for failure/A/module-not-loaded and must
move with the relock; test_integration_logoscore.cpp has the same two
hang-masked assertions.

checks.tests green: logosctl 250 + 30 + 29, logoscore 20 + 27.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): protocol 303ab08 — the last level of the carrier chain

Bumps logos-protocol, logos-plugin-qt and logos-liblogos together. The three
levels below are logos-plugin-qt#30, logos-qt-sdk#46 and logos-liblogos#190.

Bumping ONLY this repo's own logos-protocol input does not work, and that is not
a lock-tidiness opinion -- it was measured. That bump moves 1 of 230 protocol
nodes, builds green, and ships a liblogos_protocol.dylib WITHOUT the change,
because the runtime library is staged by a CARRIER rather than by this input. Of
everything in the closure only logos-protocol, logos-plugin-qt and
logos-liblogos carry one; cpp-sdk, capability-module, package-manager,
package-downloader and test-modules do not. Nothing fails when you get this
wrong -- the lock diff is real, the build is green, and the daemon loads a
library without the fix.

The lock still holds 230 protocol nodes at 16 revs afterwards. That is the usual
explosion and it is not what ships; the closure is the claim:

    protocol paths in the built daemon closure: exactly ONE
    n7yilrs2...-logos-protocol-lib-0.8.0        (the build of 303ab08)

    shipped liblogos_protocol.dylib:
      "(any)" (UTF-16)   x1      <- present only in 303ab08
      old warning        x0
      new warning        x2      <- whenObjectAvailable + the fixed onEventWhenAvailable

Both halves of that are needed. `strings` cannot see "(any)" because
QStringLiteral is UTF-16, and the new warning text appears once even in the OLD
library because whenObjectAvailable has always used it -- checking either alone
reports the wrong answer, which it did here first time round.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(watch): both subscription forms are now one call

logos-protocol#74 let onEventWhenAvailable take an EMPTY event name as the
wildcard, so the detour the wildcard form needed is gone: whenObjectAvailable()
+ requestObject() + onEvent(), a QPointer guard and a nested lambda collapse
into the same one-line call the named form already made.

The detour existed because ONE guard rejected three unrelated arguments at once
-- an empty object name and a null callback, which are unusable, and an empty
event name, which is meaningful and which the plain onEvent has always honoured.
Routing around it kept `watch <module>` with no --event working, but left the
CLI's DEFAULT invocation on a different code path from its --event form, which
is the shape a silent regression hides in.

Requires the lock bump in the previous commit: against the old library
onEventWhenAvailable answers 0 for an empty name, so the wildcard would refuse.
That is not a latent trap -- WildcardWatchSucceedsTheInstantLoadModuleReturns
fails loudly on exactly it, in both suites.

checks.tests-logoscore and checks.tests-logosctl both green (20+27 and
250+30+29), status read from nix rather than from a pipeline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(call): mirror the loaded-set gate cases into the logoscore suite

The gate landed with coverage in tests/test_integration.cpp only, so the
`logoscore` binary had none for it -- and the two binaries compile from separate
files, so "the other suite covers it" is not true here. The `watch` cases for
the same gate are already in both.

Ports both cases plus the hardening of NoLoadNegativePaths: `timeout` also exits
non-zero, so its two "call must not succeed" assertions passed BY hanging, and
now exclude 124 explicitly.

Verified as coverage, not as compilation. With the gate deleted from
callModuleMethod and nothing else changed:

    FAILED  ErrorPathTest.NoLoadNegativePaths                         (25012 ms)
    FAILED  ErrorPathTest.CallOnAModuleThatIsNotLoadedIsRefusedNotAwaited (10929 ms)
    OK      ErrorPathTest.CallImmediatelyAfterLoadStillReachesTheModule (924 ms)

The two negative cases fail on the timeouts they exist to forbid, and the
positive control stays green -- which is the point of shipping it alongside
them: it is what a shortened acquire deadline would break while the negative
cases still passed. Restored, both suites green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 23:02:47 -03:00
Dario Lipicar 500f5de1f1 chore: bump protocol to 0.9 (#105)
* chore: bump protocol to 0.9

* chore(deps): relock liblogos host stack

* bump

* bump

* test(access-policy): use token-bound callers
2026-08-25 20:43:14 -03:00
Dario Gabriel LipicarandClaude Opus 5 162dbc9fff fix(daemon stop): stop losing the shutdown reply, and stop calling that a failure
`logosctl daemon stop` printed {"code":"RPC_FAILED","message":"shutdown RPC
call failed."} and exited 3 for shutdowns that had already succeeded. It cost
the "Stop the daemon" step of doctests/logosctl-daemon.test.yaml one failure
out of nine identical shutdowns in the same macOS CI job; the daemon really
had stopped, and `daemon status` two seconds later said so.

Two independent defects, one on each side of the call.

DAEMON. CoreServiceImpl::shutdown() returned {"status":"ok"} and left the
event loop from a detached std::thread that slept 200ms and called
QCoreApplication::quit(). The reply is not on the wire at that point: the
transport serialises it after the handler returns and hands it to the socket,
which only pushes it out when the event loop services that socket's write
notifier. quit() is not a queued event -- QCoreApplication::exit() interrupts
the dispatcher directly -- so if the main thread was descheduled for longer
than the sleep, the loop came back, exited, and the buffered reply died with
the process. QtRO surfaces no transport error for this; the client just waited
out its 20s deadline and saw nothing.

The quit now runs on the main thread, from a timer, and drains the event loop
before ending it. The 200ms is now a courtesy margin rather than the
correctness mechanism, and $LOGOSCTL_SHUTDOWN_GRACE_MS makes it settable --
including to 0, which the new regression test uses because it is the setting
that used to lose the reply outright.

QtRO offers nothing better: QRemoteObjectHostBase has no per-reply
write-completion signal and no client-disconnect signal, so "quit when the
response has actually been flushed" is not reachable without forking Qt, and
the daemon also serves plain TCP/TLS through a different transport.

CLIENT. RpcClient::shutdown() reported RPC_FAILED whenever the reply was not
an object -- including when there was no reply. But a missing reply is the
expected outcome of asking a process to die, and both docs said so already:
docs/spec.md promised "the client treats the connection loss as a successful
shutdown" and docs/project.md promised exit 0 for it. Neither was implemented.

It now answers the question the reply was standing in for, from evidence: the
pid recorded in daemon/state.json (snapshotted before the call, since a clean
shutdown deletes that file) is watched for up to 15s, or for a remote daemon
the endpoint is re-probed. Gone means success, with `confirmed_by` naming the
evidence; still running means a real error, with a message that says which.
Blindly treating silence as success would have been the more dangerous
mistake -- a wedged daemon is also silent -- so it is not what this does.

That inference is only sound about a pid that was alive to begin with, so
`stop` now refuses a stale session up front the way `daemon status` already
does: a state.json naming this client's instance and a dead pid means there is
no daemon to stop (NO_DAEMON, exit 2). Without it, a session left behind by
last week's daemon would "connect" to nothing, time out, observe that the pid
is gone, and call that a successful shutdown.

TESTS.
  * ShutdownReplyTest.StopSucceedsWithNoGracePeriod (integration): 60
    start/stop cycles at LOGOSCTL_SHUTDOWN_GRACE_MS=0, asserting the command
    succeeds, the daemon is actually gone, and the reply arrived rather than
    being reconstructed from the process exiting. Measured through this
    fixture on macOS: 6 losses in 100 cycles before the daemon fix, 0 in 120
    after.
  * CommandTest.Stop_StaleSession_* : the stale-session guard, its live-pid
    control, and the remote-client case it must not block. CommandTest now
    isolates HOME and the config dir, so the suite no longer reads whichever
    ~/.logosctl the developer happens to have.
  * ProcessUtil.WaitForProcessExit* : the primitive the confirmation rests on.

A/B over the shipped binaries, 30 stop cycles per arm at zero grace, macOS:
pre-fix 8 failures; daemon fix only 0 (no reply lost); client fix only 0
(20 replies lost, every command still correct); both 0. At the default 200ms
grace both arms are clean, which is why this presented as a rare CI flake.

Independent of PR #99: that PR does not touch either function, and the two
diffs do not overlap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 18:08:04 -03:00
Dario Gabriel LipicarandClaude Opus 5 ed19258375 fix(core_service): report METHOD_FAILED from the error channel, not a null value
callModuleMethod judged failure with `ret.is_null()` because it called the
one invokeRemoteMethod overload that has no CallError* parameter
(logos_api_client.h:352-356, whose body forwards to the QVariant overload
with the error channel dropped). A method that legitimately returns null was
therefore indistinguishable from a call that failed — and that single line
was the entire empirical basis for the qt-generator's refusal to allow an
optional return.

Switching to the CallError-carrying overload (logos_api_client.h:98) is not
sufficient on its own: an unknown method name is deliberately NOT reportable
on the wire (logos_protocol.h:274-279 says so outright, and the cdylib
dispatch ends `return nullptr;  // unknown method`), so a naive !err.ok()
would have turned every typo into a silent success. The decision is now:

  !err.ok()                            -> METHOD_FAILED + {code,message,origin}
  result is a dispatch_failed envelope -> METHOD_FAILED (the provider refused)
  null AND method provably not exposed -> METHOD_NOT_FOUND + available_methods
  otherwise                            -> ok, null included

METHOD_NOT_FOUND is not invented — docs/spec.md:918 specified that envelope,
with available_methods, all along; core_service simply never produced it. It
costs one extra round-trip only on a null return.

The logic lives in a new pure unit, core_service/call_envelope.{h,cpp}, with
no Qt and no logos-protocol, which is what makes it unit-testable at all. The
value path is byte-identical: the same nlohmannArgsToQVariantList /
qvariantToNlohmann the json overload used internally.

Behaviour changes a reviewer must agree with: a null return is now `ok`
rather than METHOD_FAILED, and a dispatch_failed envelope returned as data is
now METHOD_FAILED rather than `ok`. No existing test encoded the old
behaviour; no exit code or ok/error verdict flipped in any fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 17:08:28 -03:00
Dario LipicarandClaude Opus 5 0b2afed18a chore(deps): track master for protocol, cpp-sdk and plugin-qt (#92)
* feat(access-policy): --access-policy enforce, and prove it on a real daemon

`--access-policy` already reached the runtime; what was missing was a way to
ask for deny-by-default without hand-writing JSON, and any evidence that it
works. The README actively said the opposite ("enforcement is not yet
implemented ... a no-op for now") — it has been enforced for a while.

resolveAccessPolicyArg moves out of main.cpp into daemon/access_policy_arg.
so it can be unit-tested, and gains one spelling: the literal `enforce`
expands to {"version":1,"mode":"enforce","restrictions":{}}. That is not a
second switch — `mode` is still the runtime's only switch — it is the bare
document that arms it. Checked before the file branch, so arming enforcement
can't depend on the daemon's working directory.

The integration tests are the point: same binaries, same modules, same call,
policy the only variable. test_ipc_module declares test_basic_module and
test_extlib_module; test_basic_module declares nothing.
  no flag  -> requestModule(test_basic_module, test_extlib_module) mints
  enforce  -> the same call is refused, and both names appear in the log
  enforce  -> requestModule(test_ipc_module, test_basic_module) still mints
The third is the one that matters; a change that refused everything would
pass the second on its own. The refusal is matched structurally rather than
by exact text because the two capability_module implementations in this tree
quote the names differently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(qt-host): link the Qt host runtime from logos-plugin-qt, not logos-qt-sdk

logoscore's daemon and its in-process core service are built on LogosAPI,
LogosAPIProvider and LogosProviderObject. Those moved out of logos-qt-sdk
into logos-plugin-qt, which publishes them as the `logos-qt-host` package
with the CMake target logos-qt-host::logos_qt_host. Point at that target.

Those three headers were the ONLY thing this repo took from logos-qt-sdk —
it emits no Qt consumer wrappers, ships no UI plugin, and never touches
logos_qt_lp_bridge.h or logos_ui_plugin_context.h — so the logos-qt-sdk
input is dropped outright rather than kept alongside. LOGOS_QT_SDK_ROOT
becomes LOGOS_QT_HOST_ROOT in all three derivations (build, tests,
buildPortable), and `--version` now reports the logos-plugin-qt commit.

Both new failure modes are hard errors, never silent skips: an unset
LOGOS_QT_HOST_ROOT is a FATAL_ERROR before find_package runs, and a
find_package that somehow does not define the imported target is a
FATAL_ERROR too.

logos-qt-host needs TokenManager::forIdentity/isolateIdentity, which
logos-protocol only grew on its per-client-token-store commit, so the
lock moves there. logos-plugin-qt is rev-pinned for now because
nix/qt-host.nix does not exist on its default branch yet.

Verified on aarch64-darwin: `nix build .#checks.aarch64-darwin.tests`
passes 21/21 with the committed lock and no overrides (same 21 as the
pre-change baseline), .#cli and .#cli-bundle-dir build, and the set of
LogosAPI/LogosAPIProvider/LogosProviderObject/qtArgDecode symbols in the
logoscore binary is identical to the pre-change build.

* chore(deps): re-pin the SDK stack onto the pushed b4 revs

Rebased onto master, so the inputs have to name the revs the rest of the b4
stack was actually pushed at rather than each input's default branch:

  logos-cpp-sdk           a04b2788  b3 codegen tip; a strict descendant of
                                    cpp-sdk master, so forward-only
  logos-protocol          c8bab12   per-client token store — logos-qt-host
                                    calls TokenManager::forIdentity, which
                                    exists nowhere else
  logos-plugin-qt         cc24fa1   was 8ccb1fc. The superset branch that
                                    logos-liblogos and logos-module-builder
                                    also pin, so exactly ONE logos-qt-host
                                    is in the closure — this CLI links it
                                    directly AND through liblogos_core
  logos-liblogos          f2a15ef   the liblogos built on that same qt-host
  logos-capability-module 0cb33fb   master, pinned explicitly — see below

All five are rev-pinned in the URL rather than left to the lock: every one is
a branch commit, so an unpinned url lets `nix flake update` silently relock
onto a default branch that does not build here.

capability_module deliberately does NOT move to the universal port (07dba1f).
That port declares metadata.json#host_services and fails closed until a host
calls logos_module_grant_host_services — and nothing in this stack calls it
yet (neither logos-liblogos nor logos-plugin-qt contains a single call site).
Built against it, the daemon's capability gate refuses EVERY requestModule
with "not granted the token_registry host service", so no module can call
another; the new access-policy integration test caught exactly that. 0cb33fb
is what logos-liblogos and logos-standalone-app lock too.

Verified on aarch64-darwin with the committed lock and no overrides:
  .#checks.aarch64-darwin.tests-logosctl   191 + 25 + 21 tests, all PASSED
  .#checks.aarch64-darwin.tests-logoscore   20 + 24 tests, all PASSED
  .#checks.aarch64-darwin.tests             built (exit 0)
  .#packages.aarch64-darwin.{cli,ctl}       built (exit 0)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): rev-pin logos-test-modules at the b4 qt-host tip

The daemon-backed integration checks load these plugins into the daemon
this repo builds, so the two share one host runtime in one process image
-- the same constraint that already rev-pins logos-liblogos. a639b934
links the test modules against logos-qt-host rather than logos-qt-sdk and
carries the matching B4 stack pins; the previous lock sat on master
(f8077fab), which predates that repoint.

The URL had to change, not just the lock. The input was an UNPINNED url,
so it resolved to the default branch -- and f8077fab IS master's tip.
`nix flake update logos-test-modules` was therefore a silent no-op that
would leave the ten b4 commits behind while reporting success.

f8077fab is a strict ancestor of a639b934 (verified on a non-shallow
clone), so this is forward-only, not a lineage switch.

Two behaviour changes ride along and were checked against this repo's
assertions rather than assumed safe:
  * test_basic_module and test_extlib_module migrate to
    interface "universal". Neither declares metadata.json#host_services,
    so the fail-closed gate that keeps logos-capability-module pinned off
    its universal port does not apply here.
  * stringLength now answers in CHARACTERS, not bytes. Every assertion
    here is ASCII ("abcdef" -> 6), so the two agree.

The access-policy fixture still has its pair: test_ipc_module declares
[test_basic_module, test_extlib_module] and test_basic_module declares
none, so basic -> extlib stays undeclared.

Checks built by name, all exit 0: tests-logosctl, tests-logoscore,
tests. 281 tests, 0 failures, 0 skips.

* test: use test_ipc_new_api_module as the transitive-dependency fixture

These integration tests pick a module that DECLARES the other two, so one
load-module has to pull all three, and then request a token across that edge.
test_ipc_module was that fixture; it is being retired as a duplicate. Its
successor declares exactly the same dependency pair, so the fixture role
transfers unchanged.

Worth doing in the same breath as the retirement rather than after: these call
GTEST_SKIP() when the module is missing, so deleting the module out from under
them would not have turned anything red — the dependency-resolution and
token-request coverage would simply have stopped running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(windows): refuse an unknown target instead of silently skipping it

`logos_use_shared_runtime_from_dll` empties the static archive of each named
IMPORTED target so the symbol resolves to liblogos_core.dll's exported copy
instead. It skipped any name that was not a target, which makes a typo or a
moved target silent — and the failure it hides is the duplicate-statics class:
the image keeps its own static copy of the shared runtime alongside the DLL's,
and PE has no interposition to collapse the two.

That hazard was already WRITTEN DOWN at basecamp's call site ("naming the old
target here would be a silent no-op … Windows would regress to the 29
'rejecting unauthorized call' lines this shim exists to prevent") — documented,
but not enforced. This enforces it.

Taken from feat/sdk-codegen-phase-a, which hardened its logoscore-cli copy and
never fixed basecamp's; feat/sdk-codegen-b3 has neither. It is the one place
where reconciling onto b3 would otherwise lose work, so both copies get it.

Behaviour is unchanged for every current caller: the function early-returns off
Windows, and both call sites pass the same two targets
(logos-qt-host::logos_qt_host, logos-protocol::logos_protocol) that phase-a's
hardened copy already accepts. x86_64-windows still evaluates (386 packages).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: use logos-co/setup-nix-cache-action for Nix setup and caching

Replaces the per-repo installer + cachix pair with the shared action, which
installs Nix with the Logos Attic cache (cache.nix.logos.co) preconfigured and
publishes what the job builds — master to the public cache, every other ref to
ci.

Each converted job also gains

    environment: ${{ github.ref == 'refs/heads/master' && 'public-cache' || '' }}

because ATTIC_TOKEN_PUBLIC only exists inside that environment. Without it the
secret resolves empty on master and publishing is silently skipped — the job
still passes, so the omission would not show up as a failure.

The action installs Nix itself on every runner, macOS included. That is a
deliberate reversal of the workaround these files carried: the comments here
said cachix/install-nix-action collides with the runner's pre-existing _nixbld
users (eDSRecordAlreadyExists), so DeterminateSystems' installer was used
instead. It no longer reproduces — logos-delivery-module has already been
converted the plain way and its `build-and-test (macos-latest)` leg passes.
Keeping the workaround would have meant a second installer plus a duplicated
substituter/key block in ten files, guarding against something two green runs
say does not happen. If it ever recurs it fails loudly at install, which is
recoverable; the silent-skip above is the failure mode worth engineering
against.

One property is deliberately NOT carried over: the old cachix step ran with
`continue-on-error: true` so a failed cache push could not fail a job whose
tests passed. The action exposes no equivalent, and adding one here would also
swallow genuine setup failures now that the same step installs Nix rather than
only publishing at the end.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: drop references to removed generator flags and interfaces

README and docs described module authoring in terms of LogosProviderBase,
LOGOS_METHOD and --provider-header, none of which exist. Updated to the
universal model, keeping the retired shapes named as history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): track master for protocol, cpp-sdk and plugin-qt

logos-protocol#59, logos-cpp-sdk#138 and logos-plugin-qt#19 merged, so the three
rev pins bridging to them are retired, each with its rationale rewritten to name
the PR that closed the gap.

Left pinned: logos-liblogos, logos-capability-module and logos-test-modules —
their branches are still in flight and no merged upstream was confirmed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:47:31 -03:00
e48fc7fef3 Add logosctl alongside logoscore: one CLI for daemon, modules and packages (#76)
* feat(core_service): add refreshModules and cascade unload by default

Two runtime prerequisites for the logosctl merge.

refreshModules() wraps logos_core_refresh_modules(), which liblogos
documents as "call after installing new modules so they become
discoverable". Basecamp calls it on the package_manager install event,
which is why installing a module there needs no restart. core_service
did not expose it, so a CLI that installs a package had no way to make
the daemon see it short of a restart.

unloadModule() now takes withDependents and the CLI defaults it to true
(--no-dependents opts out). logos_core_unload_module already accepted
the flag; core_service hardcoded false, which left dependents running
against an unloaded provider. The result now carries dependents_unloaded
so the cascade is reported rather than silent.

The dispatch entry defaults a missing second argument to true, so a
one-argument unloadModule call keeps working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(flake): bundle package_manager and package_downloader

logoscore bundled only capability_module, so it could authenticate but
not manage packages — that was lgpd's and lgpm's job, as separate
binaries. Bundling the two package modules is what lets one binary do
the whole job.

Same trio logos-basecamp bundles, assembled the same way (map the
install bundler over the module libs), so the CLI and the GUI drive an
identical module surface rather than the CLI being a reduced sibling.
Only the package manager ships a distinct lib-portable; the other two
are variant-agnostic, matching basecamp's split.

Verified against a real daemon: all three are discovered with no module
configuration, both package modules load, and package_downloader
resolves the live default catalog.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): make the config dir a self-contained session

The daemon knew about ~/.logoscore only as a place to keep its own
state files; packages, trust material and persistence lived elsewhere
or nowhere. Now the config dir is the whole world for a session:

  <configDir>/modules   installed core modules (writable)
  <configDir>/plugins   installed UI plugins
  <configDir>/keyring   trusted signing keys
  <configDir>/cache     downloaded .lgx
  <configDir>/data      module persistence

so copying the directory carries the session's packages and its trust
assumptions with it, and two sessions can disagree about both.

<configDir>/modules joins the search path beside the bundled dir.
Without it an installed module would sit on disk that the daemon could
never see, and install-then-load could not work at all.

The bundled package modules are loaded at boot and pointed at these
directories -- the same four setters basecamp calls -- because every
package command is an RPC into them. All best-effort: a daemon that
cannot manage packages is still fully usable for loading and calling
modules, so none of it aborts startup.

Verified live: a bare daemon creates the tree, loads all three modules,
and reports the embedded packages via getInstalledPackages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(package): daemon-side install, upgrade and remove

Adds the mutating package operations the CLI never had, orchestrated
inside the daemon and exposed as core_service.planPackageOperation /
applyPackageOperation, plus the `package`, `catalog` and `key` command
groups on top.

Why daemon-side: package_manager gates destructive work behind a
listener-ack protocol with a 3-second deadline. Driving that from a
short-lived client would mean holding an event subscription open,
interleaving it with outbound calls, and winning a three-second race
across the RPC boundary. In-process the ack cannot lose that race, and
every client command stays thin and stateless.

plan/apply is split so `--dry-run` and the confirmation prompt see
exactly what apply will do -- the same dependency-change table basecamp
shows, including which running modules get stopped. Without -y and
without a TTY the operation is refused rather than assumed-yes, so a
script that forgot --yes fails loudly instead of silently uninstalling.

install/upgrade take dependencies, remove takes dependents, both by
default. Installing never loads: it puts files on disk, and only
modules already running beforehand are restarted afterwards.

Verified against the live catalog on a portable build: install
openmetrics; install chat_module pulling delivery_module in order;
re-install as a no-op; install then load with no daemon restart
(refreshModules); and removing delivery_module cascading through
chat_module with both stopped first.

One trap worth naming: LogosList{vec} does not wrap a std::vector the
way it wraps a scalar -- it yields an empty args array, and the module
sees a zero-argument call it cannot dispatch. The batch uninstall
builds its argument explicitly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(config): replace the flag surface with a YAML session document

Configuration was ~20 flags plus two hand-written mini-grammars: a
`NAME=PROTOCOL[,k=v...]` parser that existed only to squeeze a nested
structure through a flag, and a per-flag defaults<config<CLI merge. Both
are gone. main.cpp drops from 943 to 531 lines.

Configuration now lives in the session:

  <configDir>/daemon/config.yaml   written by `daemon config set`
  <configDir>/client/config.yaml   written by `client config set`

and is never passed alongside an unrelated command, so `daemon start`
and every client command take the session exactly as it is on disk.
--config-dir is the one surviving flag, because it selects *which*
session to act on and so cannot itself live inside one.

The split is by audience: files a human edits are YAML, files the
daemon and modules own stay JSON (state.json, tokens, the auto token).
Converting through nlohmann::json means the existing validated
daemonConfigFromJson / clientStateFromJson keep doing the schema work.

Two traps fixed while wiring it up, both of the accept-then-ignore kind
that leaves an operator with no explanation:

  - A bare `modules: {core_service: [ ... ]}` sequence was silently
    skipped (only the `{transports: [...]}` spelling parsed). It is now
    accepted as shorthand.
  - Unknown top-level keys are rejected by `config set` and the error
    names the correct spelling, so `insecureTcp` no longer looks like it
    worked when the key is `insecure_tcp`.

Module search paths remain configurable via the `modules_dirs` key,
which is what replaces -m for tests and dev loops.

The eight CLI tests that covered deleted flags are rewritten against
the new surface: malformed YAML rejected without clobbering the
existing config, unknown keys named, set/show round-trip, and absent
config treated as defaults rather than an error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(cli): logosctl, with docker-style command groups

Renames the binary and reorganises ~20 flat hyphenated commands into
groups: daemon, client, module, token, package, catalog, key, plus
top-level aliases for the four verbs that cannot be confused with a
runtime module (status, call, watch, stats) and the two package verbs
with no module meaning (install, search).

The hyphenated names survive as internal dispatch tokens but are hidden
from --help: `module load X` is rewritten to `load-module X` in argv
before CLI11 parses. The rewrite happens in argv rather than via nested
CLI11 subcommands because daemonSub->fallthrough() pushes a nested
subcommand's unmatched arguments up to the top level, where they are
rejected ("The following argument was not expected: show").

`module` is no longer an alias for the verbose call syntax -- it is the
group. Use `call`.

Also implements --detach, which was specified but missing. It re-execs
rather than continuing in the forked child: macOS refuses to let a
process that has already initialised CoreFoundation keep running after
fork(), and the Qt/liblogos link pulls CoreFoundation in before main.
The child redirects stdio to <configDir>/daemon/daemon.log -- without
that the shell never sees EOF and `daemon start --detach` appears to
hang -- and the parent returns only once state.json exists, so the next
command cannot race the boot.

Env vars and the default session directory rename to LOGOSCTL_*
and ~/.logosctl. User-facing messages now name the group grammar rather
than the internal tokens.

Verified on the portable build: daemon start --detach returns in ~3s
with a working daemon; catalog ls, search, install --dry-run, install,
package ls, module ls/load/show, upgrade (no-op), and remove of a
loaded module all behave. 18/18 CLI tests, unit tests unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: rewrite for logosctl, sessions and YAML config

The README documented a flag surface that no longer exists
(--persist-config, --module-transport, --modules-dir, the seven
--client-* flags) and had no account of sessions or packages at all.

Replaces the daemon/transport/persist-config sections with: what a
session directory is and why it is portable, `daemon config set` and
the YAML schema, and the package/catalog/keyring commands. Keeps the
two hard-won warnings that are still true -- a remote daemon must expose
capability_module as well as core_service, and plaintext tcp on a
non-loopback host needs an explicit opt-in.

Doctests are renamed and rewritten around sessions: the daemon spec no
longer passes -m but seeds ./session/modules, and uses
`daemon start --detach` instead of backgrounding with & (which returned
before the transports bound and raced the first command).

Also fixes the stats table: the MODULE column was a fixed 12 characters,
so a real name like "test_basic_module" ran straight into the PID with
no separator.

Verified the rewritten daemon-doctest sequence by hand against a dev
build: seed session, start --detach, module ls/load, call, stats, stop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* doctests: use --detach and the session log

`daemon start --detach` already returns only once the daemon is
accepting commands and sends its output to <session>/daemon/daemon.log,
so the `sh -c '... > logs.txt 2>&1 &'` wrapper is not just redundant --
it hid the output the specs then tried to cat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: ship logosctl alongside logoscore instead of replacing it

logosctl is new and unvalidated; logoscore is what people depend on
today. Replacing one with the other in a single step meant every
consumer had to move at once, on trust. Shipping both means logosctl
can be validated in real use first, and logoscore removed afterwards.

Both binaries build from this repo over one shared runtime -- daemon,
core_service, client, output. They differ only in main.cpp and a
Config::Flavor the front-end sets, which selects the config directory,
the env var consulted for an override, the config file names, and the
format they are written in.

The isolation is the point, so it is deliberate and tested:

  logoscore  ~/.logoscore   LOGOSCORE_CONFIG_DIR  daemon/config.json
  logosctl   ~/.logosctl    LOGOSCTL_CONFIG_DIR   daemon/config.yaml

A logosctl session cannot disturb a logoscore deployment. Reading needs
no branch -- YAML is a superset of JSON, so one parser handles both --
only writing differs.

logoscore is behaviourally unchanged, which took two specific
decisions:

  - The session directory and the package-module bootstrap are gated on
    the modern flavor. Auto-loading two extra modules would change what
    `status` and `list-modules` report, and logoscore's doc-tests assert
    those exact counts.
  - The bundled package modules live in modules-pkg/ rather than
    modules/, because logoscore scans the latter and would otherwise
    report two modules it never had.

Verified: `logoscore --help` is the old flat surface with all four flag
families intact; a logoscore daemon reports loaded:1 not_loaded:0 and
creates no session directories; both daemons run at once with separate
state.

Its doc-tests are restored unchanged. logosctl gets its own, including
a new logosctl-packages spec covering the capability that motivated the
merge -- search, dry-run, install, load, remove -- verified end to end
against the live catalog.

122/129 unit tests, 18/18 CLI tests. The 7 OutputTest failures are
pre-existing on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* build: give logosctl its own flake outputs

Building both binaries into one package meant `nix build` and `.#cli`
started handing out logosctl too, which is the opposite of keeping the
two apart while the new one is validated.

Now each output ships exactly one binary:

  .#cli  .#cli-bundle-dir  .#cli-appimage   ->  logoscore
  .#ctl  .#ctl-bundle-dir  .#ctl-appimage   ->  logosctl
  .#     (default)                          ->  logoscore

So anything already pointing at the default or at `.#cli` -- including
every doc-test across the workspace that does
`nix build github:logos-co/logos-logoscore-cli` -- keeps getting the
tool it gets today, and logosctl is strictly opt-in.

They still compile together, since they share everything but main.cpp;
only the packaging is split. modules-pkg/ ships solely in the ctl
outputs, because logoscore never scans it.

logoscore's desktop entry and icon are restored, and logosctl gets its
own. The logosctl doc-tests now build .#ctl / .#ctl-bundle-dir.

Verified: every output builds and ships only its own binary; both
portable bundles run and report the module set expected of each.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(config): let each session subdirectory be redirected

The session directory being self-contained is what makes it portable, and
that should stay the default -- but it was also the only option, which
made reasonable setups impossible: sharing one keyring across sessions,
putting the .lgx cache on a bigger disk, or pointing at a modules tree
something else manages.

A `dirs:` block now redirects any of them:

    dirs:
      keyring: ~/.config/logos/trusted-keys
      cache: /var/cache/logos
      modules: /opt/logos/modules
      plugins: plugins-custom
      data: /var/lib/logos/data

The form of the value decides whether portability survives, which is the
part worth knowing:

    plugins-custom  -> <session>/plugins-custom   still portable
    ~/x             -> $HOME/x                    outside the session
    /var/cache/...  -> as given                   outside the session

`~` is handled because it is the natural thing to write in a config file
and would otherwise resolve to <session>/~/... , which exists nowhere.

Overrides resolve once, when set, so relocating a session afterwards
cannot silently drag an absolute path along with it. Only the daemon
applies them, and it does so before anything asks Config for a path.

persistence_path is folded into dirs.data -- it was already the same
setting under an older name -- so the two no longer need choosing
between at the point of use.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): rotating log file with configurable size and retention

--detach used to dup2 stdout/stderr straight onto daemon/daemon.log,
which grew without bound and had no rotation. A long-lived daemon needs
better than that.

There is now a logs/ directory, like Basecamp's, and a logging block:

    logging:
      enabled: true          # false -> no log file at all
      file: daemon.log       # inside dirs.logs
      max_size_mb: 10        # rotate past this; 0 = never rotate
      max_files: 5           # keep this many in total
      console: true          # mirror to the terminal

dirs.logs joins the overridable session directories, so logs can be
shipped somewhere a collector already watches.

Capture is pipe-based rather than a file redirect, and that is the
load-bearing decision: module hosts are separate processes holding
inherited descriptors. Redirecting to a file catches their output but
makes rotation impossible -- renaming a file out from under a child that
has it open just keeps filling the old inode. A pipe puts one reader in
charge, so rotation is safe and subprocess output still lands in the
log. Same shape as basecamp's LogRedirector, which solved this already.

The size cap and retention come from spdlog's rotating sink rather than
being hand-rolled; liblogos already logs through spdlog. Lines arriving
from the pipe already carry their own timestamp and level, so the sink
uses a raw pattern instead of stamping them twice.

Two bugs found while testing it:

  - Draining raced shutdown. stop() cleared the running flag before
    restoring the descriptors, so a reader holding data would process
    it, loop, see the flag clear and exit -- dropping whatever was still
    in the pipe. The last lines before a shutdown are exactly the ones
    worth keeping. EOF is now the only stop condition.
  - --detach reported the wrong path. The parent prints before the child
    has read the config, so it guessed the default and lied to anyone
    who had redirected dirs.logs or renamed the file. It now reads the
    same config the child will.

Verified live: default, disabled, and redirected-with-custom-filename
all behave and are reported accurately. Four unit tests cover capture
of both streams, no double-stamping, rotation with retention, and
disabled-is-not-an-error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): timestamp log files, and bound the directory

Adopts basecamp's naming -- each start writes its own
daemon_<yyyymmdd_HHMMSS>.log -- so a session's output is one file you can
point at, instead of every run appending into the same daemon.log.

Two things beyond copying basecamp:

  - `logging.file` survives as a symlink to whichever file is current, so
    `tail -F logs/daemon.log` follows across restarts and nobody has to
    work out a stamp. It also means --detach can report a path that is
    always valid; previously it had to guess one, and guessed wrong for
    anyone who had redirected dirs.logs.

  - max_files now bounds the *directory*, pruning oldest-first at each
    start. spdlog's retention only prunes within one sink's rotation set,
    and every start opens a new stamped base name, so without this a
    daemon restarted a hundred times would leave a hundred logs behind.
    basecamp has exactly that problem.

Verified live: three restarts leave three stamped files with the symlink
tracking the newest; five restarts with max_files: 2 leave two.

Two new tests cover the naming and the symlink resolving to the current
session, and the cross-session pruning. The rotation test needed fixing
too -- it counted the symlink as a log file, which predated the symlink
existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: ignore suffixed nix out-links

.gitignore listed `result` but not `result-*`, so every out-link from a
targeted build -- `nix build '.#ctl' -o result-ctl`, `-o result-tests`,
and so on -- was untracked-but-not-ignored, and `git add -A` committed
them as symlinks into /nix/store.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: keep releasing logoscore, and release logosctl beside it

The earlier rename left the release workflow building the `cli-*`
outputs -- which are logoscore -- while naming every artifact
`logosctl-*`. A release/** push would have shipped logoscore binaries
under the wrong name, and stopped releasing logoscore under its own.

Both are now built and published as separate, correctly-named assets:

  logoscore-{x86_64,aarch64}-linux.tar.gz   from .#cli-appimage
  logoscore-aarch64-macos.tar.gz            from .#cli-bundle-dir
  logosctl-{x86_64,aarch64}-linux.tar.gz    from .#ctl-appimage
  logosctl-aarch64-macos.tar.gz             from .#ctl-bundle-dir

logoscore's asset names are exactly what they were, which matters:
release sets fetch this repo and expect a bundle containing
`bin/logoscore`. Each tool builds from its own flake outputs, so an
asset labelled logoscore contains logoscore and nothing else.

Both jobs gained a tool matrix with fail-fast disabled, so a failure in
the under-validation logosctl cannot block a logoscore release. The
release job now collects artifacts by pattern instead of naming each
one, so retiring logoscore later means deleting a matrix entry rather
than unpicking a download list.

Release notes lead with logoscore as the tool to use, and say the two
share no state so installing logosctl cannot disturb an existing setup.

The doc-tests workflow globs doctests/*.test.yaml, which now covers both
suites, so it is no longer named after one of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: run both tools' suites, in parallel

While both binaries ship, both get tested. logoscore had no automated
coverage on this branch at all -- only its doc-tests -- so a change to
the shared runtime could regress the tool people actually use and
nothing would say so.

tests/test_cli_logoscore.cpp and tests/test_integration_logoscore.cpp
are copies of the suites frozen against logoscore's surface. Copies
rather than a parameterised shared suite on purpose: the two surfaces
genuinely differ, and this way retiring logoscore is a delete rather
than an unpick.

checks.tests-logosctl and checks.tests-logoscore are separate
derivations, so nix builds them concurrently; checks.tests aggregates
both, keeping `nix build .#checks.<sys>.tests` working for CI while now
covering both tools.

It immediately earned its keep, catching three regressions:

  - The integration harness still passed -m, which logosctl no longer
    accepts, so its daemon never started and seven integration tests
    were failing on this branch. It now writes the modules_dirs config
    the daemon reads.

  - `logoscore --version` reported "logosctl version ...". The version
    banner had been renamed wholesale; each front-end now names itself.
    Exactly the sort of thing nobody notices until a bug report cites
    the wrong tool.

  - The new log sink only mirrored to the console when stdout was a
    TTY, so `logoscore -D > logs.txt` -- which the doc-tests do --
    produced an empty file. Mirroring now follows the configured
    setting, pipe or terminal alike, and the log file is gated to
    logosctl so logoscore's output behaviour is untouched.

Both suites green: logosctl 138 unit + 18 CLI + 18 integration,
logoscore 20 CLI + 18 integration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(package): honour -o, and stop parsing command lines backwards

`package download -o DIR` accepted the flag and threw it away -- the
argument was parsed into a variable and then explicitly discarded with
`(void)outDir;`. The file went to $TMPDIR regardless. The config's
`dirs.cache` had the same problem from the other end: the directory was
created and documented as holding downloads, and nothing ever wrote to
it.

The cause was the same for both. package_downloader takes no
destination, so the file lands in $TMPDIR on the DAEMON's filesystem --
which is where the move has to happen too. Doing it client-side would
work only for a local daemon. So `downloadPackage` joins the daemon-side
package operations: it downloads, then moves the result into the
requested directory, or into the session's cache/downloads when no -o
was given. The client resolves a relative -o against its own working
directory first, so a local daemon does what the user typed; against a
remote one the path is remote, and a bad one fails loudly rather than
quietly writing elsewhere.

Writing the first test for it turned up something worse. CLI11's
`parse(std::vector<std::string>&)` consumes the vector from the BACK --
only the rvalue overload reverses for you -- so passing natural order
parses the command line backwards. `watch` and `issue-token` did reverse
first; nothing else did. It goes unnoticed with one positional and flags
(order does not matter), and is quietly wrong the moment an option takes
a value, because the option pairs with the token to its LEFT:

  package download pkg -o dir   ->  name="dir", output="pkg"
  package install a b --version 1.0
                                ->  names=["1.0","b"], version="a"

So `package install`, `search --category`, and `download -o` all
misparsed. Every site now goes through one `parseArgs` helper that
reverses, which fixes the broken ones, is a no-op for the harmless ones,
and removes the trap for the next command.

PackageCommand had no unit tests at all, which is why a discarded flag
survived review. Four now cover download; the two asserting -o reaches
the daemon fail against the old code.

142 unit + 18 CLI + 18 integration green for logosctl, 20 + 18 for
logoscore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: one README about the repo, one document per tool

The README had grown into a logosctl manual with a banner on top telling
logoscore users that everything below did not apply to them, and pointing
them at doc-test YAML for their actual documentation. Since logoscore is
still the tool to use, its documentation should not be the thing you are
told to skip.

So: README.md covers what is true of both -- what the repo is, the two
binaries and how they differ, the flake outputs, the test targets,
dependency resolution, platforms -- and hands off to one document per
tool.

  docs/logoscore.md   the usage material, unchanged, as its own document
  docs/logosctl.md    sessions, config, logs, packages, examples

Writing logosctl's own document exposed a gap: it had no command
reference at all. The rewrite dropped the client-command list, argument
typing and exit codes, and left behind a "see Argument typing below"
pointing at a section that no longer existed. All three are back, with
the command list written against the grammar that is actually
implemented (checked against normalizeGroupVerbs and the subcommand
dispatch, not from memory), plus the two defaults worth stating up front
-- install does not load, remove takes dependents.

Also fixes stale copy that survived the earlier rewrite: `load-module`
where logosctl says `module load`, and a "multiple module directories"
caption over a --config-dir example, from a flag logosctl does not have.

Deleting logoscore later is now deleting one file and a table row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(daemon): make TLS configurable again, and say why startup failed

Three bugs, all found by running the doc-tests I had just rewritten
instead of trusting them.

**tcp_ssl could not be configured at all.** `transportFromJson` never
read `cert` or `key`. That was harmless while those arrived via
`--module-transport ...,cert=...,key=...`, parsed by the CLI mini-grammar
-- but that grammar is gone, and the config file is now the only place to
set them. So every tcp_ssl listener bound with no certificate: the daemon
started, reported itself healthy, accepted connections, and failed every
handshake with "no shared cipher (SSL routines)". The client just saw
"core_service not reachable".

The stripping was deliberate but applied one layer too high: cert and key
have no business in state.json, which clients read, but the config file is
where an operator *authors* them. `transportToJson` now takes
`includeSecrets` -- true writing the config, false writing state.json. A
test asserts the round-trip, and another asserts the key path never
appears in state.json.

**`--detach` swallowed the reason startup failed.** Config validation runs
before LogSink opens the log, and the child's stderr went to /dev/null, so
a rejected config produced "daemon exited during startup. See
<path>/logs/daemon.log" -- naming a file that had never been created. The
child's early output now goes to a startup file the parent reads and
prints on failure, removed either way. LogSink takes those descriptors
over as soon as it starts, so the file only ever holds pre-logging output.

**The plaintext-TCP guard advertised a flag that does not exist.** It said
"pass --insecure-tcp"; logosctl has no such flag. It now names the config
key, `insecure_tcp: true`.

Verified end to end against a real daemon: plaintext guard refuses and
says why, loopback TCP binds and serves `status`/`module ls` from a
separate client session, TLS serves the same over 6443/6444, and dropping
the CA while keeping verify_peer still fails closed.

144 unit + 18 CLI + 18 integration green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(flake): give autoPatchelf the libraries both binaries now link

Every Linux build failed:

  auto-patchelf could not satisfy dependency libyaml-cpp.so.0.8
  wanted by .../bin/.logoscore-wrapped

The packaging derivations listed only Qt in buildInputs, which is what
autoPatchelfHook resolves DT_NEEDED entries against. yaml_json.cpp and the
log sink are in the shared sources, so *both* binaries link yaml-cpp and
spdlog -- including logoscore, which is why its Linux build broke too on a
branch that was supposed to leave it alone.

macOS does not patchelf, so this was invisible locally and in the macOS
CI jobs; only the Linux matrix caught it, and it took down the AppImage
builds, the CI job, and every Linux doc-test with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(doctests): bring the logosctl specs up to what logosctl does

Nineteen doc-test steps were failing. None of them were runtime bugs in
the specs' own right -- they were specs still describing an older
logosctl, which is its own kind of failure: a doc-test that lies is worse
than no doc-test.

  transports      Still drove `--module-transport` and hand-written
                  client/config.json. The flags had been dropped from the
                  `run:` lines but no config step replaced them, so the
                  daemon never bound TCP at all and every step after it
                  failed. Rewritten around `daemon config set` /
                  `client config set` with YAML documents, for both the
                  plaintext and TLS halves.

  daemon          Read the log at session/daemon/daemon.log; logs moved to
                  session/logs/. The crash-recovery step passed `-m`,
                  which logosctl does not accept, so its daemon never
                  started and the step reported LEAKED against a worker
                  that had never existed.

  modules-bundle  Asserted all three modules in result/modules. The
                  package modules live in modules-pkg/ so that logoscore's
                  modules/ stays byte-identical -- which the spec is now
                  the place that explains.

  packages        Expected the interactive wording ("dry run",
                  "Installed:"). Doc-tests are not a terminal, so every
                  command renders JSON. The install was working the whole
                  time; only the assertions were wrong. They now match the
                  JSON, and the prose says why it is JSON.

Rewriting the transports spec is what turned up the TLS and --detach bugs
fixed in the previous commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): a typo must not abort the daemon, and a key must not lie

Three defects in the YAML config path, all found by building a Python
client against this CLI and checking its assumptions against the binary
rather than the docs.

**A config typo aborted the process.**

  printf 'version: 2\nmodules_dirs: /single/path\n' > bad.yaml
  logosctl --config-dir ./s daemon config set ./bad.yaml
  => libc++abi: terminating due to uncaught exception ...
     [json.exception.type_error.302] type must be array, but is string

nlohmann's `json::value(key, default)` THROWS when the key is present
with the wrong type, every config read used it, and nothing caught it.
So it was not one key -- it was every key in both readers. A scalar
where a list belongs is an ordinary mistake and it killed the binary.

Now a type-checked reader (src/json_schema.h) records
"<dotted.path>: expected <what>, but got <what>" and the document is
refused whole, the same shape as the existing unknown-key error:

  {"code":"INVALID_CONFIG",
   "message":"modules_dirs: expected a list of strings, but got a string."}

Both readers went through it, including two paths that could abort the
daemon mid-boot rather than at `config set`.

**`config set` validated after writing.** A schema-invalid document was
installed and then reported as an error, leaving the session holding a
config the daemon would refuse to boot from. Validation now happens
entirely in memory first, on both the daemon and client sides -- the
client side had no schema validation at all -- and the write is
temp-file + rename instead of truncate-in-place.

That exposed a fourth: `yaml_json::dump` emitted numeric-looking strings
bare, so `port: "6001"` came back as the number 6001. The bytes
validated were not the bytes written.

**Two keys were accepted, stored, and never applied.**

`signature_policy` sat on the allowlist and was written verbatim to
config.yaml but was never even parsed. An operator setting `require` got
no enforcement and no warning. It is now parsed with a strict allowlist
and pushed into package_manager at boot beside setKeyringDirectory --
the module has had setSignaturePolicy all along. Unset issues no RPC, so
the module keeps its own default instead of having it restated.

The top-level `ssl: {cert, key, ca}` block was parsed into DaemonConfig
and read by nobody; only per-listener cert/key reached the transport
set. Configuring TLS the obvious way therefore produced listeners with
no certificate and "no shared cipher" on every handshake -- the same
failure fixed one layer down last commit. It is now a session-wide
default that per-listener values override.

Also: docs advertised `module load --no-deps`, which does not exist --
`module load` takes only a positional name and always resolves
dependencies. Corrected, along with the rest of the command reference,
verified against the binary.

logoscore is untouched: 20 CLI + 18 integration, exactly as before.
logosctl 171 unit (was 144) + 25 CLI (was 18) + 18 integration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(detach): re-exec the launcher, not the ELF it hides

`daemon start --detach` was dead on Linux portable builds. The daemon
exited immediately with status 127, no output, and no log file -- so the
only diagnostic was "daemon exited during startup. See <path>", naming a
file that had never been created.

strace, on a real Linux box, said it in one line:

  execve(".../bin/.logosctl.elf", [...]) = -1 ENOENT
  exit_group(127)

A portable bundle installs the CLI as a launcher script beside a hidden
companion:

  bin/logosctl        the launcher, a shell script
  bin/.logosctl.elf   the real ELF

The launcher exists because that ELF cannot be started on its own: its
PT_INTERP names a dynamic loader that is not on the host, so the launcher
runs it through a known-good ld.so instead. The ENOENT is the kernel
reporting the missing *interpreter* -- the ELF is right there.

--detach re-execs itself (it has to: macOS forbids running a forked
process that has initialized CoreFoundation), and it re-exec'd
executablePath(), which is that ELF.

My first attempt preferred argv[0], reasoning that it is what the caller
actually typed. That was wrong, and the trace showed it failing
identically: the launcher execs ld.so with the ELF, ld.so drops itself
from argv, and the program sees the ELF as argv[0] too. Neither source
of truth names the launcher.

So the mapping is applied to whatever candidate we end up with, using the
convention the launcher script itself documents -- the install dir is the
one holding the hidden companion `.$BASE.elf`. `bin/.logosctl.elf` maps
back to `bin/logosctl`. argv[0] is still preferred over
executablePath() (it is what was invoked, and it is right when a bare
name resolves through PATH), and it is absolutised, since the daemon may
run from a different directory.

Only this combination was ever broken: portable AND Linux AND --detach.
macOS bundles a real binary with qt.conf and no launcher, Linux dev
builds are ordinary ELFs, and the foreground -D path never re-execs. The
one doc-test that uses the portable bundle is the packages spec, and
cachix served a permanent 522 for one of its store paths from the day it
was written -- so its 14 cascading failures read as infrastructure until
the cache recovered and the real failure surfaced underneath.

Verified on Linux against the same bundle that failed: daemon starts
detached, all three bundled modules load, `daemon stop` returns ok.

Also here, and what made the diagnosis possible: --detach now prints the
TAIL of the daemon log rather than its path. The startup file only holds
output from before LogSink takes the descriptors, so a daemon that dies
after logging is up left it empty and the reason unread. That there was
no log at all is what pointed at exec.

179 unit tests (8 new, covering the launcher mapping and its edges: no
sibling, an ordinary foo.elf, a non-executable candidate, absent argv[0]).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* build: bundle logosctl/logoscore as headless Qt programs

* build: bump nix-bundle-dir and nix-bundle-appimage to main

Picks up the merged trampoline drop: per-arch psABI PT_INTERP, DT_RPATH,
qtCliApp for headless Qt, and the AppImage consumer that already tracks
the same pin. nix-bundle-dir 4fd87d1 (PR tip) → cb9afc8; appimage
8fcc56b → 04a3cf8.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-03 21:49:15 -03:00
Dario Lipicar 5623313431 fix: reap stale socket files at daemon boot (#71)
The reported symptom was dozens of lingering /tmp/logos_* files. Two
separate paths produce them, and only one was fixed:

  * A graceful stop now unlinks its own sockets — the SIGTERM self-pipe
    handler lets QCoreApplication::exec() return, so QLocalServer's
    destructor (the only thing that unlinks) actually runs. That landed
    in logos-module-loader-qt#4 and is what eliminated the bulk of the
    volume, since previously *every* shutdown leaked one file per module.

  * Nothing cleans up after a hard kill. SIGKILL, a module crash, and the
    PR_SET_PDEATHSIG kill that reaps orphaned logos_host children all
    skip destructors by construction, so those files accumulate forever.

logos::reapStaleSockets has been available (and gtested) in logos-protocol
since #20 but had no production caller. Wire it into daemon boot.

Placed after the already-running check and before logos_core_init: at
that point none of our sockets exist yet, so the reaper cannot race its
own endpoints. Co-resident nodes are safe because logos::isSocketDead
fails closed — it unlinks only an S_ISSOCK inode that we own and whose
connect() is refused — so a live socket and a regular file sharing the
prefix are both left alone. QDir::tempPath() is the authoritative
directory: QLocalServer resolves bare server names against it, and it
honours $TMPDIR on Linux and macOS alike.

Two integration tests, in a private $TMPDIR so the assertions describe
this node rather than a dev box's leftovers:

  NoSocketsSurviveGracefulStop — loads a module (so the tally includes a
  logos_host child socket), asserts sockets exist, stops, asserts none
  remain. The "sockets exist" precondition is asserted so the test cannot
  pass vacuously.

  BootReapsStaleSocketsButSparesLiveOnesAndFiles — plants a stale socket
  (bound, listener closed), a live socket held open by the test process,
  and a regular file named logos_execution_zone-1.0.0.lgx, then boots a
  daemon and asserts only the stale one is gone. The last two are the
  fail-closed guarantees: reaping the live socket would break a
  co-resident node, and a glob-based cleanup would delete a
  multi-hundred-MB build artefact that really does sit in the temp dir.

Verified by negative control: with the reapStaleSockets call disabled,
BootReapsStaleSockets... FAILS and the rest stay green. Full suite 18/18
on aarch64-darwin.

Socket paths are capped by sockaddr_un::sun_path (104 bytes on macOS),
and macOS's default temp dir is already ~50 chars, so the fixture picks a
directory short enough to hold "/logos_capability_module_<12 hex>" and
skips if none fits — otherwise every listen() fails with a confusing
HostNotFoundError.
2026-07-22 11:34:55 -03:00
Dario LipicarandClaude Opus 4.8 679a9af8fd fix: report module version in list-modules and module-info (#59) (#60)
* fix: report module version in list-modules and module-info (#59)

The version column was always empty and the JSON omitted version entirely
(`delivery_modulev` in the table = name + "v" + empty). The data layer
never populated it.

Source it generically from liblogos' new logos_core_get_modules_info(),
which returns name/path/loaded/dependencies/dependents/metadata per known
module. listModules and getModuleInfo now build from that single call, so:

- list-modules shows VERSION (table) and "version" (JSON) for loaded AND
  not_loaded modules (version comes from on-disk metadata).
- module-info reports version plus dependencies/dependents.
- load-module / reload-module responses include the version.

Output: empty versions render as "-" (modules aren't required to declare
one), and the NAME/VERSION table columns size to their content so long
names no longer collide with the version.

Tests: OutputTest cases for the table/dash/collision rendering; an
integration ReportsModuleVersion test (real daemon) covering version +
dependencies across list-modules/module-info/load-module; doctest
assertions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat: report uptime for loaded modules (#59)

list-modules, status, and module-info now report uptime_seconds for loaded
modules, derived from the load timestamp liblogos records (now - loaded_at).
Unloaded modules report no uptime_seconds (uptime is loaded-only). The
daemon stamps loaded_at with the same wall clock core_service reads, so the
value is consistent.

Tests: ReportsModuleVersion asserts uptime_seconds is absent for an
unloaded module and present once loaded.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: bump logos-liblogos to merged modules-info API (#159)

Re-pin logos-liblogos 87ae7ce → 819faac (master, #159) and its transitive
logos-module a3e288a → 2ec64c4 (#21), so the version/uptime/deps features
build against the merged generic modules-info API without overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat: add --no-json/--human flag; doctest shows both output forms

Every client command auto-selects human output on a TTY and JSON when piped.
Add --no-json (alias --human) as the explicit inverse of --json, forcing the
human-readable form even when piped — useful for scripts/log capture and for
docs that want to show the terminal view deterministically.

The "Running Modules with the logoscore Daemon" doctest now shows both the
human table and the JSON for status, list-modules, module-info, stats, and
call, using --human/--json to render each form.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(doctest): show human and JSON output samples for both forms

The doctest generator renders commands and prose but not captured output,
so add post_text blocks displaying both the human-readable and JSON output
for status, list-modules, module-info, stats, and call. The steps already
run both forms (--human/--json) and assert on them; this surfaces the
outputs in the generated tutorial.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(doctest): restore unrelated generated outputs

run.sh clears outputs/ wholesale and only regenerates the spec(s) passed to
it, so regenerating just logoscore-daemon.md inadvertently dropped the
transports and concurrent-blocking tutorials. Restore them unchanged — this
PR only touches the daemon doctest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 20:46:22 -03:00
Iuri Matias bfb64580b0 fix: add core_service token 2026-06-05 17:16:46 -04:00
Iuri MatiasandCursor c410a01091 replace QJson* with nlohmann/json
replace QJson* with nlohmann/json

update flake

ci fix

update flake

update flake

fix include path to use SDK's include/cpp where logos_api_client.h actually lives

The combined SDK symlinkJoin package places headers at include/cpp/ (from
the headers sub-package), not at include/ root. Using include/cpp ensures
the compiler finds the SDK's logos_api_client.h (with nlohmann::json overloads)
before the stale copy shipped inside logos-liblogos's include directory.

Co-authored-by: Cursor <cursoragent@cursor.com>

fix watch command args order for CLI11 parse

CLI11's parse(vector<string>) processes from the back of the vector, so
args must be reversed before calling parse(). The 0532673 commit removed
this reversal, causing the positional module arg and --event value to be
swapped when both are present.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 15:39:58 -04:00
Iuri Matias 1ba40efead ci fix 2026-05-21 16:21:36 -04:00
Iuri Matias ab7b2179bf replace QString 2026-05-21 15:28:27 -04:00
Dario Lipicar 823fa438bc test module subprocess crash scenario (#29)
* test module subprocess crash scenario

* bump liblogos
2026-05-20 11:56:38 -03:00
Dario Lipicar 454e0696e9 add integration tests (#27)
* add integration tests

* PR comments
2026-05-15 10:56:11 -03:00