Commit Graph
5 Commits
Author SHA1 Message Date
Dario LipicarandClaude Opus 5 0b2afed18a chore(deps): track master for protocol, cpp-sdk and plugin-qt (#92)
* feat(access-policy): --access-policy enforce, and prove it on a real daemon

`--access-policy` already reached the runtime; what was missing was a way to
ask for deny-by-default without hand-writing JSON, and any evidence that it
works. The README actively said the opposite ("enforcement is not yet
implemented ... a no-op for now") — it has been enforced for a while.

resolveAccessPolicyArg moves out of main.cpp into daemon/access_policy_arg.
so it can be unit-tested, and gains one spelling: the literal `enforce`
expands to {"version":1,"mode":"enforce","restrictions":{}}. That is not a
second switch — `mode` is still the runtime's only switch — it is the bare
document that arms it. Checked before the file branch, so arming enforcement
can't depend on the daemon's working directory.

The integration tests are the point: same binaries, same modules, same call,
policy the only variable. test_ipc_module declares test_basic_module and
test_extlib_module; test_basic_module declares nothing.
  no flag  -> requestModule(test_basic_module, test_extlib_module) mints
  enforce  -> the same call is refused, and both names appear in the log
  enforce  -> requestModule(test_ipc_module, test_basic_module) still mints
The third is the one that matters; a change that refused everything would
pass the second on its own. The refusal is matched structurally rather than
by exact text because the two capability_module implementations in this tree
quote the names differently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(qt-host): link the Qt host runtime from logos-plugin-qt, not logos-qt-sdk

logoscore's daemon and its in-process core service are built on LogosAPI,
LogosAPIProvider and LogosProviderObject. Those moved out of logos-qt-sdk
into logos-plugin-qt, which publishes them as the `logos-qt-host` package
with the CMake target logos-qt-host::logos_qt_host. Point at that target.

Those three headers were the ONLY thing this repo took from logos-qt-sdk —
it emits no Qt consumer wrappers, ships no UI plugin, and never touches
logos_qt_lp_bridge.h or logos_ui_plugin_context.h — so the logos-qt-sdk
input is dropped outright rather than kept alongside. LOGOS_QT_SDK_ROOT
becomes LOGOS_QT_HOST_ROOT in all three derivations (build, tests,
buildPortable), and `--version` now reports the logos-plugin-qt commit.

Both new failure modes are hard errors, never silent skips: an unset
LOGOS_QT_HOST_ROOT is a FATAL_ERROR before find_package runs, and a
find_package that somehow does not define the imported target is a
FATAL_ERROR too.

logos-qt-host needs TokenManager::forIdentity/isolateIdentity, which
logos-protocol only grew on its per-client-token-store commit, so the
lock moves there. logos-plugin-qt is rev-pinned for now because
nix/qt-host.nix does not exist on its default branch yet.

Verified on aarch64-darwin: `nix build .#checks.aarch64-darwin.tests`
passes 21/21 with the committed lock and no overrides (same 21 as the
pre-change baseline), .#cli and .#cli-bundle-dir build, and the set of
LogosAPI/LogosAPIProvider/LogosProviderObject/qtArgDecode symbols in the
logoscore binary is identical to the pre-change build.

* chore(deps): re-pin the SDK stack onto the pushed b4 revs

Rebased onto master, so the inputs have to name the revs the rest of the b4
stack was actually pushed at rather than each input's default branch:

  logos-cpp-sdk           a04b2788  b3 codegen tip; a strict descendant of
                                    cpp-sdk master, so forward-only
  logos-protocol          c8bab12   per-client token store — logos-qt-host
                                    calls TokenManager::forIdentity, which
                                    exists nowhere else
  logos-plugin-qt         cc24fa1   was 8ccb1fc. The superset branch that
                                    logos-liblogos and logos-module-builder
                                    also pin, so exactly ONE logos-qt-host
                                    is in the closure — this CLI links it
                                    directly AND through liblogos_core
  logos-liblogos          f2a15ef   the liblogos built on that same qt-host
  logos-capability-module 0cb33fb   master, pinned explicitly — see below

All five are rev-pinned in the URL rather than left to the lock: every one is
a branch commit, so an unpinned url lets `nix flake update` silently relock
onto a default branch that does not build here.

capability_module deliberately does NOT move to the universal port (07dba1f).
That port declares metadata.json#host_services and fails closed until a host
calls logos_module_grant_host_services — and nothing in this stack calls it
yet (neither logos-liblogos nor logos-plugin-qt contains a single call site).
Built against it, the daemon's capability gate refuses EVERY requestModule
with "not granted the token_registry host service", so no module can call
another; the new access-policy integration test caught exactly that. 0cb33fb
is what logos-liblogos and logos-standalone-app lock too.

Verified on aarch64-darwin with the committed lock and no overrides:
  .#checks.aarch64-darwin.tests-logosctl   191 + 25 + 21 tests, all PASSED
  .#checks.aarch64-darwin.tests-logoscore   20 + 24 tests, all PASSED
  .#checks.aarch64-darwin.tests             built (exit 0)
  .#packages.aarch64-darwin.{cli,ctl}       built (exit 0)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): rev-pin logos-test-modules at the b4 qt-host tip

The daemon-backed integration checks load these plugins into the daemon
this repo builds, so the two share one host runtime in one process image
-- the same constraint that already rev-pins logos-liblogos. a639b934
links the test modules against logos-qt-host rather than logos-qt-sdk and
carries the matching B4 stack pins; the previous lock sat on master
(f8077fab), which predates that repoint.

The URL had to change, not just the lock. The input was an UNPINNED url,
so it resolved to the default branch -- and f8077fab IS master's tip.
`nix flake update logos-test-modules` was therefore a silent no-op that
would leave the ten b4 commits behind while reporting success.

f8077fab is a strict ancestor of a639b934 (verified on a non-shallow
clone), so this is forward-only, not a lineage switch.

Two behaviour changes ride along and were checked against this repo's
assertions rather than assumed safe:
  * test_basic_module and test_extlib_module migrate to
    interface "universal". Neither declares metadata.json#host_services,
    so the fail-closed gate that keeps logos-capability-module pinned off
    its universal port does not apply here.
  * stringLength now answers in CHARACTERS, not bytes. Every assertion
    here is ASCII ("abcdef" -> 6), so the two agree.

The access-policy fixture still has its pair: test_ipc_module declares
[test_basic_module, test_extlib_module] and test_basic_module declares
none, so basic -> extlib stays undeclared.

Checks built by name, all exit 0: tests-logosctl, tests-logoscore,
tests. 281 tests, 0 failures, 0 skips.

* test: use test_ipc_new_api_module as the transitive-dependency fixture

These integration tests pick a module that DECLARES the other two, so one
load-module has to pull all three, and then request a token across that edge.
test_ipc_module was that fixture; it is being retired as a duplicate. Its
successor declares exactly the same dependency pair, so the fixture role
transfers unchanged.

Worth doing in the same breath as the retirement rather than after: these call
GTEST_SKIP() when the module is missing, so deleting the module out from under
them would not have turned anything red — the dependency-resolution and
token-request coverage would simply have stopped running.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(windows): refuse an unknown target instead of silently skipping it

`logos_use_shared_runtime_from_dll` empties the static archive of each named
IMPORTED target so the symbol resolves to liblogos_core.dll's exported copy
instead. It skipped any name that was not a target, which makes a typo or a
moved target silent — and the failure it hides is the duplicate-statics class:
the image keeps its own static copy of the shared runtime alongside the DLL's,
and PE has no interposition to collapse the two.

That hazard was already WRITTEN DOWN at basecamp's call site ("naming the old
target here would be a silent no-op … Windows would regress to the 29
'rejecting unauthorized call' lines this shim exists to prevent") — documented,
but not enforced. This enforces it.

Taken from feat/sdk-codegen-phase-a, which hardened its logoscore-cli copy and
never fixed basecamp's; feat/sdk-codegen-b3 has neither. It is the one place
where reconciling onto b3 would otherwise lose work, so both copies get it.

Behaviour is unchanged for every current caller: the function early-returns off
Windows, and both call sites pass the same two targets
(logos-qt-host::logos_qt_host, logos-protocol::logos_protocol) that phase-a's
hardened copy already accepts. x86_64-windows still evaluates (386 packages).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: use logos-co/setup-nix-cache-action for Nix setup and caching

Replaces the per-repo installer + cachix pair with the shared action, which
installs Nix with the Logos Attic cache (cache.nix.logos.co) preconfigured and
publishes what the job builds — master to the public cache, every other ref to
ci.

Each converted job also gains

    environment: ${{ github.ref == 'refs/heads/master' && 'public-cache' || '' }}

because ATTIC_TOKEN_PUBLIC only exists inside that environment. Without it the
secret resolves empty on master and publishing is silently skipped — the job
still passes, so the omission would not show up as a failure.

The action installs Nix itself on every runner, macOS included. That is a
deliberate reversal of the workaround these files carried: the comments here
said cachix/install-nix-action collides with the runner's pre-existing _nixbld
users (eDSRecordAlreadyExists), so DeterminateSystems' installer was used
instead. It no longer reproduces — logos-delivery-module has already been
converted the plain way and its `build-and-test (macos-latest)` leg passes.
Keeping the workaround would have meant a second installer plus a duplicated
substituter/key block in ten files, guarding against something two green runs
say does not happen. If it ever recurs it fails loudly at install, which is
recoverable; the silent-skip above is the failure mode worth engineering
against.

One property is deliberately NOT carried over: the old cachix step ran with
`continue-on-error: true` so a failed cache push could not fail a job whose
tests passed. The action exposes no equivalent, and adding one here would also
swallow genuine setup failures now that the same step installs Nix rather than
only publishing at the end.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: drop references to removed generator flags and interfaces

README and docs described module authoring in terms of LogosProviderBase,
LOGOS_METHOD and --provider-header, none of which exist. Updated to the
universal model, keeping the retired shapes named as history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): track master for protocol, cpp-sdk and plugin-qt

logos-protocol#59, logos-cpp-sdk#138 and logos-plugin-qt#19 merged, so the three
rev pins bridging to them are retired, each with its rationale rewritten to name
the PR that closed the gap.

Left pinned: logos-liblogos, logos-capability-module and logos-test-modules —
their branches are still in flight and no merged upstream was confirmed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:47:31 -03:00
Dario Lipicar 2d6fba858e fix(logoscore): make --persistence-path reach the directory the daemon reads (#85)
* fix(logoscore): make --persistence-path reach the directory the daemon reads

The flag parsed and was then dropped. The CLI merge in main_legacy.cpp wrote
only DaemonConfig::persistencePath, while Daemon::start redirects the session
data directory from cfg.dirs.data. The alias that folds persistence_path into
dirs.data lives inside daemonConfigFromJson, so it only ever ran while READING
a config file -- never after the CLI merge. With no config.json the flag was
ignored; with one, the FILE beat the command line.

Write both fields on an explicit --persistence-path, so this block has the
precedence it already documents: explicit command line > config file > default.
The config-file rule (dirs.data beats persistence_path) is unchanged.

The value is absolutised first: a dirs: entry in config.json is relative to the
config dir, but a path typed on the command line is relative to the cwd, which
is what this flag has always meant and what --modules-dir still means.

Drop the now-dead 'const auto& persistencePath = cfg.persistencePath' binding
in Daemon::start -- the fossil of the regression, unused since the switch to
dirs.data.

Regression tests in tests/test_integration_logoscore.cpp assert the flag
CHANGES BEHAVIOUR rather than that it appears in --help: they boot a real
daemon, load test_basic_module, and check where logos_core mkpath'd the
instance directory, for all three reachable states (no config.json, config.json
with dirs.data, config.json with only persistence_path) plus a relative path.

* fix(logoscore): let a `~/` --persistence-path expand instead of absolutising it

Follow-up to the fix in this branch, from its own verification round.

Routing the CLI value through `dirs.data` puts it through `resolveOverride`
(src/config.cpp), which does TWO things to a non-absolute value: it expands a
leading `~/` against $HOME, and otherwise reparents the path under the config
dir.  The absolutise added here to defeat the second one also defeated the
first: `std::filesystem::absolute("~/x")` is `$PWD/~/x`, which `resolveOverride`
then passes through untouched because it is already absolute.  The result is a
real directory literally named `~` beside the cwd.

That also left the two spellings of the same string disagreeing -- a config.json
`dirs.data: "~/x"` resolves to `$HOME/x` while `--persistence-path '~/x'` did
not -- which is the opposite of what writing both fields is for.

So skip the absolutise for a leading `~/` and let resolveOverride do its job.
Every other relative form still pins to the shell's cwd, which is what this flag
has always meant.

The shell expands an unquoted `~/x` before the binary ever sees it, so this
governs only a quoted or script-built value -- exactly the case where the user
cannot have meant a directory called `~`.

Pinned by a new integration test, TildeFlagExpandsAgainstHome, which boots a
real daemon with `--persistence-path '~/tilde-data'` in the fixture's isolated
HOME and asserts the module's instance directory lands under $HOME, that no
literal `~` directory appears beside the cwd, and that the default data dir
stays empty.

NOT COMPILED LOCALLY: `nix develop` in this repo rebuilds logos-protocol /
logos-cpp-sdk / logos-qt-sdk from the pinned revisions, none of which are in the
local store -- the same wall the verification round hit.  CI is the gate for
this one.  Verified by reading instead: resolveOverride's tilde branch exists on
origin/master too (config.cpp:81-84, `std::getenv("HOME")`), the fixture sets
HOME to an isolated directory (test_integration_logoscore.cpp:154), and the new
test uses only helpers already defined in the file.
2026-08-12 11:48:59 -03:00
Dario LipicarandClaude Opus 5 3ed617e167 feat(windows): cross-compile logosctl.exe, and run a module on Windows (#84)
* feat(windows): resolve the executable path and PATH on Windows

paths.cpp had __APPLE__ and __linux__ branches and no fallback, so on Windows
executablePath()/executableDir() returned "" SILENTLY. Everything layered on
them then failed for no visible reason: bundledModulesDir(),
bundledPackageModulesDir() and relaunchPath() all start from the executable's
directory, so module discovery and daemon relaunch would simply find nothing.

  - executablePath()/executableDir() get a GetModuleFileNameW branch. The
    helper grows its buffer rather than trusting MAX_PATH: the API truncates
    instead of failing when the buffer is too small (and on older Windows does
    not even NUL-terminate).
  - access(p, X_OK) is not expressible on Windows -- there is no execute
    permission bit and no X_OK -- so the mode becomes F_OK. Callers already
    pair this with fs::is_regular_file(), which is the meaningful test.
  - PATH is ';'-separated, not ':'. Without this, the "invoked by bare name"
    branch of relaunchPath() would treat the whole of PATH as one directory.

The portable-bundle ".elf launcher" mapping is left alone: it is specific to
the Linux bundle layout and simply never matches on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(windows): one correct process-liveness check for both callers

daemon.cpp and status_command.cpp both asked "is this pid still running?" via
kill(pid, 0), with OPPOSITE polarity:

    daemon.cpp          alive => refuse to start, a node is already up
    status_command.cpp  dead  => report the state file as stale

so a wrong answer breaks in two different directions -- "always alive" makes
the daemon unrestartable after a crash, "always dead" lets two daemons run and
clobber each other's state.json. That is why this is now one shared helper
rather than two open-coded checks.

The POSIX contract being reproduced is kill(pid, 0): 0 = exists, EPERM =
exists but owned by someone else (still ALIVE), ESRCH = dead. Only ESRCH means
dead, which the old status_command check already got right and which the
helper preserves.

Windows equivalent uses OpenProcess + WaitForSingleObject(h, 0), NOT
GetExitCodeProcess: a process handle is signalled exactly when the process has
exited, whereas GetExitCodeProcess returns STILL_ACTIVE (259) which is
indistinguishable from a process that genuinely exited with code 259.
ERROR_ACCESS_DENIED from OpenProcess maps to the EPERM case -- the pid exists,
it is simply not ours -- and so counts as alive.

Verified natively: self => alive, an absent pid => dead, pid 1 (root-owned,
EPERM) => alive, pid 0 => dead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(windows): cross-compile logosctl.exe, and run a module on Windows

logosctl.exe now builds for x86_64-w64-mingw32, starts its daemon on real
Windows and loads capability_module through QPluginLoader:

  logosctl daemon start --detach  -> Daemon started (pid 13528)
  logosctl status                 -> "modules":[{"name":"capability_module",
                                      "status":"loaded"}], "loaded":1

New src/platform_compat.h holds the shims needed in more than one place.
Three are traps rather than translations:
  * gmtime_s/localtime_s take the SAME arguments as gmtime_r in the OPPOSITE
    order, so this is a swap, not the rename the compiler suggests.
  * timegm has no counterpart; _mkgmtime is the equivalent. mktime would
    silently shift every timestamp by the local UTC offset.
  * strptime has no Windows counterpart AT ALL. The hand-rolled replacement is
    deliberately strict and %n-anchored because its one caller treats a parse
    failure as "expired" -- a lax parser would keep a bad token alive.

config.cpp: %LOCALAPPDATA%, then %USERPROFILE%, then %TEMP%. Non-roaming on
purpose -- pids, sockets, ports and logs are machine-specific. The absolute-path
test needs BOTH `raw[0] == '/'` and fs::path::is_absolute(), because
fs::path("/x").is_absolute() is FALSE on Windows.

log_sink.cpp: _pipe plus the SetStdHandle half. dup2 rebinds only the CRT fd
table, while a CreateProcess child inherits STARTUPINFO std handles from
STD_OUTPUT_HANDLE -- without SetStdHandle the module-host output this sink
exists to capture escapes entirely.

port_allocator.cpp: Winsock, and SO_EXCLUSIVEADDRUSE rather than SO_REUSEADDR.
On Windows SO_REUSEADDR permits binding a port that is already bound AND
LISTENING, so the probe could have handed out a port another process was
serving on and the daemon would have advertised it. Also a socket_t/
socketValid abstraction: SOCKET is unsigned, so the old `fd < 0` check is never
true and a failed socket() would have sailed into bind().

daemon.cpp: the POSIX self-pipe is kept verbatim; Windows uses
SetConsoleCtrlHandler and posts quit() via a queued invokeMethod, since the OS
runs that routine on its own thread. Deliberately NO self-pipe there:
QEventDispatcherWin32 routes QSocketNotifier to WSAAsyncSelect, which needs a
real SOCKET, and a CRT pipe fd fails WSAENOTSOCK while the dispatcher discards
the return value -- the notifier would simply never fire.

main.cpp: --detach becomes one CreateProcess with DETACHED_PROCESS |
CREATE_NEW_PROCESS_GROUP, implemented rather than stubbed out. Readiness polls
WaitForSingleObject, not GetExitCodeProcess, whose STILL_ACTIVE (259) cannot be
told apart from a genuine exit code 259.

CMakeLists: imported logos_core needs IMPORTED_IMPLIB and IMPORTED_LOCATION set
separately -- getting it wrong is only a CMake WARNING, and it substitutes
logos_core-NOTFOUND into the link line so the build dies much later. The test
block is skipped wholesale for a Windows host, not just its targets, so the
gtest FetchContent fallback cannot try to download inside the sandbox.

$out has no lib/: PE has no rpath, so a library in lib/ is one nothing can
find. All 24 non-system DLLs sit in bin/ beside the executables, verified by
resolving every import table against that directory.

KNOWN GAP: package_manager does not load on Windows yet, and says so loudly
("Package commands will be unavailable in this session") rather than failing
quietly.

* fix(windows): take the shared runtime from liblogos_core.dll

logosctl and logoscore link liblogos_core.dll AND, through
logos-qt-sdk::logos_qt_sdk, the liblogos_qt_sdk.a / liblogos_protocol.a static
archives. Once liblogos_core.dll became the single provider of the shared C++
runtime, both sides defined it and the Windows link failed outright:

    multiple definition of `TokenManager::instance()'
      liblogos_protocol.a(token_manager.cpp.obj)
      liblogos_core.dll.a(liblogos_core_dll_d000294.o): first defined here

plus LogosAPIClient::invokeRemoteMethod, LogosProviderObject::callMethodStdBridge
/ getMethodsStdBridge / setEventListenerStdBridge, and logos::reapStaleSockets.

Same fix Basecamp already carries: compile with LOGOS_SHARED_USE_DLL so the
references become __declspec(dllimport), and replace the static archives with an
EMPTY one by repointing IMPORTED_LOCATION.

The empty archive is the non-obvious half and cannot be replaced by the export
macro alone: ld picks archive members by OBJECT FILE, not by symbol, so one
object referencing something unrelated (measured in the sibling repo: a
std::string move constructor) drags LogosAPI, LogosAPIClient and TokenManager in
behind it. No export list prevents that; the archive must not be a candidate.
Repointing IMPORTED_LOCATION rather than dropping target_link_libraries keeps
the imported targets' include dirs and Qt/OpenSSL/nlohmann usage requirements.

Verified: builds clean (0 multiple-definition errors) and produces a real
PE32+ x86-64 logosctl.exe. The single-provider property is confirmed on the
STORAGE symbol, not the function symbol -- logosctl.exe carries an auto-import
thunk for TokenManager::instance (one TU references it without the dllimport
declaration), which looks like a definition in the symbol table but is not:
    logosctl.exe  _ZZN12TokenManager8instanceEvE8instance = 0 slots
    main_ui.dll   0 slots (known-good consumer)
    logos_host.exe 2 slots (correct -- separate process, no liblogos_core link)
Zero storage slots means no second instance; the thunk jumps into the DLL.

Found only because priming cachix built all 11 Windows targets. My earlier
validation built the two NEW CI targets and inferred the rest -- building what
you changed is not building what CONSUMES what you changed. The other 8 targets
were unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(deps): re-pin the chain to its merged revs

Levels 1-8 of the Windows chain are merged.  This branch was locked to
pre-merge revs of every one of them, so it could only evaluate against the
unmerged branches.

Level 9; the workspace re-pin follows once this lands.

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:32:47 -03:00
Dario LipicarandClaude Opus 5 d015a40c5f Install deadline, honest failure reasons, and a -v that actually reaches the daemon (#83)
* fix(package): give installs a real deadline, and stop eating the reason

Two defects, found from a crashed install on a real server.

**A 20-second deadline on an operation that downloads a blockchain node.**
`Impl::invoke` never passed a timeout, so every call used the
transport's default `Timeout()` -- 20 seconds. That is right for "is the
daemon up" and useless for "fetch and install blockchain_module". The
client gave up at 20s and reported

    Error: applyPackageOperation('install') RPC call failed.

while the daemon carried on and finished the job. The package landed on
disk and the user was told it had failed. Reproduced on the server: the
install completes, the client does not wait for it.

The comment above that call already said "the default RPC deadline is
far too short for that". It sat above a call that passed no deadline at
all.

planPackageOperation now gets 2 minutes (it reads the catalog) and
applyPackageOperation / downloadPackage 30 minutes (they transfer and
install). Generous on purpose: waiting too long costs a slow command,
waiting too little costs telling someone their install failed while it
is still running and about to succeed.

**The failure message threw away the reason.** package_ops returns two
error shapes -- its own step failures carry `failed_step` + `error`,
while anything that fails before the chain starts (an unreachable
module, a dead daemon, a transport error) carries `code` + `message`.
The client rendered only the first, so the second printed as

    Error: install failed at step '?':

with nothing after the colon. That is what a real diagnosis looked like
when the daemon died mid-install: it had said why, and we dropped it.

Now whichever shape arrives is reported, the step is omitted rather than
printed as '?' when unknown, and an error carrying no reason at all says
so instead of trailing off.

182 unit tests, 3 new covering both shapes and the empty case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(daemon): let -v reach spdlog so forwarded module logs survive

* fix(daemon): -v was being passed as persistConfig

`Daemon::start` ends in two adjacent bool parameters, both defaulted:

    static int start(int argc, char* argv[],
                     const DaemonConfig& cfg,
                     const std::string& configSource,
                     bool persistConfig = false,
                     bool verbose = false);

logosctl called it with one bool:

    Daemon::start(argc, argv, mergedCfg, configSource, g_verbose);

so g_verbose became persistConfig and verbose stayed false for the
entire life of the binary. It compiled, because the defaults made the
short call legal. logoscore's own call site passes both and is correct;
this was introduced when logosctl dropped --persist-config and the
argument was removed without accounting for its position.

Two consequences, both silent. Every `if (verbose)` branch in the daemon
was dead, so -v changed nothing daemon-side -- it only ever reached
main.cpp's Qt message handler, which reads the global directly. And -v
quietly enabled config persistence, which is not a thing logosctl even
offers: its configuration is installed with `daemon config set`.

The call site now names both arguments. The declaration drops its
defaults, so the class of mistake is a build error rather than a silent
slide: two adjacent same-typed defaulted parameters are exactly the shape
where omitting one is undetectable.

This is also the real reason four earlier attempts at "module logs do not
reach the session log" measured as no-ops. The code meant to raise
spdlog's level sits behind `if (verbose)` and never ran.

182 unit + 25 CLI + 18 integration green for logosctl; logoscore
unchanged at 20 + 18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: %d for QMessageLogContext::line, which is an int

Copilot caught this on the module-loader PR, and the same mistake is
here in both front-ends: the Critical/Fatal branches print
`context.line` with %u, but QMessageLogContext::line is an int. A
signed/unsigned format mismatch is undefined behaviour and can trip
-Wformat.

logoscore's two sites are included. The rendered output is identical for
any real line number, so this is not the behaviour change its frozen
suites exist to catch -- and they stay green at 20 + 18 to show it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: bump liblogos for the module-host logging fix

Picks up logos-liblogos#173, which carries logos-module-loader-qt#5 --
the host installing its own Qt message handler so a module's
diagnostics are not diverted to journald.

This is the last link. The measurements in this PR were taken with the
loader spliced in via --override-input; with this pin they hold for a
plain `nix build` of the real chain.

Single input moved (plus its own transitive loader pin).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 21:08:08 -03:00
e48fc7fef3 Add logosctl alongside logoscore: one CLI for daemon, modules and packages (#76)
* feat(core_service): add refreshModules and cascade unload by default

Two runtime prerequisites for the logosctl merge.

refreshModules() wraps logos_core_refresh_modules(), which liblogos
documents as "call after installing new modules so they become
discoverable". Basecamp calls it on the package_manager install event,
which is why installing a module there needs no restart. core_service
did not expose it, so a CLI that installs a package had no way to make
the daemon see it short of a restart.

unloadModule() now takes withDependents and the CLI defaults it to true
(--no-dependents opts out). logos_core_unload_module already accepted
the flag; core_service hardcoded false, which left dependents running
against an unloaded provider. The result now carries dependents_unloaded
so the cascade is reported rather than silent.

The dispatch entry defaults a missing second argument to true, so a
one-argument unloadModule call keeps working.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(flake): bundle package_manager and package_downloader

logoscore bundled only capability_module, so it could authenticate but
not manage packages — that was lgpd's and lgpm's job, as separate
binaries. Bundling the two package modules is what lets one binary do
the whole job.

Same trio logos-basecamp bundles, assembled the same way (map the
install bundler over the module libs), so the CLI and the GUI drive an
identical module surface rather than the CLI being a reduced sibling.
Only the package manager ships a distinct lib-portable; the other two
are variant-agnostic, matching basecamp's split.

Verified against a real daemon: all three are discovered with no module
configuration, both package modules load, and package_downloader
resolves the live default catalog.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): make the config dir a self-contained session

The daemon knew about ~/.logoscore only as a place to keep its own
state files; packages, trust material and persistence lived elsewhere
or nowhere. Now the config dir is the whole world for a session:

  <configDir>/modules   installed core modules (writable)
  <configDir>/plugins   installed UI plugins
  <configDir>/keyring   trusted signing keys
  <configDir>/cache     downloaded .lgx
  <configDir>/data      module persistence

so copying the directory carries the session's packages and its trust
assumptions with it, and two sessions can disagree about both.

<configDir>/modules joins the search path beside the bundled dir.
Without it an installed module would sit on disk that the daemon could
never see, and install-then-load could not work at all.

The bundled package modules are loaded at boot and pointed at these
directories -- the same four setters basecamp calls -- because every
package command is an RPC into them. All best-effort: a daemon that
cannot manage packages is still fully usable for loading and calling
modules, so none of it aborts startup.

Verified live: a bare daemon creates the tree, loads all three modules,
and reports the embedded packages via getInstalledPackages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(package): daemon-side install, upgrade and remove

Adds the mutating package operations the CLI never had, orchestrated
inside the daemon and exposed as core_service.planPackageOperation /
applyPackageOperation, plus the `package`, `catalog` and `key` command
groups on top.

Why daemon-side: package_manager gates destructive work behind a
listener-ack protocol with a 3-second deadline. Driving that from a
short-lived client would mean holding an event subscription open,
interleaving it with outbound calls, and winning a three-second race
across the RPC boundary. In-process the ack cannot lose that race, and
every client command stays thin and stateless.

plan/apply is split so `--dry-run` and the confirmation prompt see
exactly what apply will do -- the same dependency-change table basecamp
shows, including which running modules get stopped. Without -y and
without a TTY the operation is refused rather than assumed-yes, so a
script that forgot --yes fails loudly instead of silently uninstalling.

install/upgrade take dependencies, remove takes dependents, both by
default. Installing never loads: it puts files on disk, and only
modules already running beforehand are restarted afterwards.

Verified against the live catalog on a portable build: install
openmetrics; install chat_module pulling delivery_module in order;
re-install as a no-op; install then load with no daemon restart
(refreshModules); and removing delivery_module cascading through
chat_module with both stopped first.

One trap worth naming: LogosList{vec} does not wrap a std::vector the
way it wraps a scalar -- it yields an empty args array, and the module
sees a zero-argument call it cannot dispatch. The batch uninstall
builds its argument explicitly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(config): replace the flag surface with a YAML session document

Configuration was ~20 flags plus two hand-written mini-grammars: a
`NAME=PROTOCOL[,k=v...]` parser that existed only to squeeze a nested
structure through a flag, and a per-flag defaults<config<CLI merge. Both
are gone. main.cpp drops from 943 to 531 lines.

Configuration now lives in the session:

  <configDir>/daemon/config.yaml   written by `daemon config set`
  <configDir>/client/config.yaml   written by `client config set`

and is never passed alongside an unrelated command, so `daemon start`
and every client command take the session exactly as it is on disk.
--config-dir is the one surviving flag, because it selects *which*
session to act on and so cannot itself live inside one.

The split is by audience: files a human edits are YAML, files the
daemon and modules own stay JSON (state.json, tokens, the auto token).
Converting through nlohmann::json means the existing validated
daemonConfigFromJson / clientStateFromJson keep doing the schema work.

Two traps fixed while wiring it up, both of the accept-then-ignore kind
that leaves an operator with no explanation:

  - A bare `modules: {core_service: [ ... ]}` sequence was silently
    skipped (only the `{transports: [...]}` spelling parsed). It is now
    accepted as shorthand.
  - Unknown top-level keys are rejected by `config set` and the error
    names the correct spelling, so `insecureTcp` no longer looks like it
    worked when the key is `insecure_tcp`.

Module search paths remain configurable via the `modules_dirs` key,
which is what replaces -m for tests and dev loops.

The eight CLI tests that covered deleted flags are rewritten against
the new surface: malformed YAML rejected without clobbering the
existing config, unknown keys named, set/show round-trip, and absent
config treated as defaults rather than an error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(cli): logosctl, with docker-style command groups

Renames the binary and reorganises ~20 flat hyphenated commands into
groups: daemon, client, module, token, package, catalog, key, plus
top-level aliases for the four verbs that cannot be confused with a
runtime module (status, call, watch, stats) and the two package verbs
with no module meaning (install, search).

The hyphenated names survive as internal dispatch tokens but are hidden
from --help: `module load X` is rewritten to `load-module X` in argv
before CLI11 parses. The rewrite happens in argv rather than via nested
CLI11 subcommands because daemonSub->fallthrough() pushes a nested
subcommand's unmatched arguments up to the top level, where they are
rejected ("The following argument was not expected: show").

`module` is no longer an alias for the verbose call syntax -- it is the
group. Use `call`.

Also implements --detach, which was specified but missing. It re-execs
rather than continuing in the forked child: macOS refuses to let a
process that has already initialised CoreFoundation keep running after
fork(), and the Qt/liblogos link pulls CoreFoundation in before main.
The child redirects stdio to <configDir>/daemon/daemon.log -- without
that the shell never sees EOF and `daemon start --detach` appears to
hang -- and the parent returns only once state.json exists, so the next
command cannot race the boot.

Env vars and the default session directory rename to LOGOSCTL_*
and ~/.logosctl. User-facing messages now name the group grammar rather
than the internal tokens.

Verified on the portable build: daemon start --detach returns in ~3s
with a working daemon; catalog ls, search, install --dry-run, install,
package ls, module ls/load/show, upgrade (no-op), and remove of a
loaded module all behave. 18/18 CLI tests, unit tests unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: rewrite for logosctl, sessions and YAML config

The README documented a flag surface that no longer exists
(--persist-config, --module-transport, --modules-dir, the seven
--client-* flags) and had no account of sessions or packages at all.

Replaces the daemon/transport/persist-config sections with: what a
session directory is and why it is portable, `daemon config set` and
the YAML schema, and the package/catalog/keyring commands. Keeps the
two hard-won warnings that are still true -- a remote daemon must expose
capability_module as well as core_service, and plaintext tcp on a
non-loopback host needs an explicit opt-in.

Doctests are renamed and rewritten around sessions: the daemon spec no
longer passes -m but seeds ./session/modules, and uses
`daemon start --detach` instead of backgrounding with & (which returned
before the transports bound and raced the first command).

Also fixes the stats table: the MODULE column was a fixed 12 characters,
so a real name like "test_basic_module" ran straight into the PID with
no separator.

Verified the rewritten daemon-doctest sequence by hand against a dev
build: seed session, start --detach, module ls/load, call, stats, stop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* doctests: use --detach and the session log

`daemon start --detach` already returns only once the daemon is
accepting commands and sends its output to <session>/daemon/daemon.log,
so the `sh -c '... > logs.txt 2>&1 &'` wrapper is not just redundant --
it hid the output the specs then tried to cat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: ship logosctl alongside logoscore instead of replacing it

logosctl is new and unvalidated; logoscore is what people depend on
today. Replacing one with the other in a single step meant every
consumer had to move at once, on trust. Shipping both means logosctl
can be validated in real use first, and logoscore removed afterwards.

Both binaries build from this repo over one shared runtime -- daemon,
core_service, client, output. They differ only in main.cpp and a
Config::Flavor the front-end sets, which selects the config directory,
the env var consulted for an override, the config file names, and the
format they are written in.

The isolation is the point, so it is deliberate and tested:

  logoscore  ~/.logoscore   LOGOSCORE_CONFIG_DIR  daemon/config.json
  logosctl   ~/.logosctl    LOGOSCTL_CONFIG_DIR   daemon/config.yaml

A logosctl session cannot disturb a logoscore deployment. Reading needs
no branch -- YAML is a superset of JSON, so one parser handles both --
only writing differs.

logoscore is behaviourally unchanged, which took two specific
decisions:

  - The session directory and the package-module bootstrap are gated on
    the modern flavor. Auto-loading two extra modules would change what
    `status` and `list-modules` report, and logoscore's doc-tests assert
    those exact counts.
  - The bundled package modules live in modules-pkg/ rather than
    modules/, because logoscore scans the latter and would otherwise
    report two modules it never had.

Verified: `logoscore --help` is the old flat surface with all four flag
families intact; a logoscore daemon reports loaded:1 not_loaded:0 and
creates no session directories; both daemons run at once with separate
state.

Its doc-tests are restored unchanged. logosctl gets its own, including
a new logosctl-packages spec covering the capability that motivated the
merge -- search, dry-run, install, load, remove -- verified end to end
against the live catalog.

122/129 unit tests, 18/18 CLI tests. The 7 OutputTest failures are
pre-existing on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* build: give logosctl its own flake outputs

Building both binaries into one package meant `nix build` and `.#cli`
started handing out logosctl too, which is the opposite of keeping the
two apart while the new one is validated.

Now each output ships exactly one binary:

  .#cli  .#cli-bundle-dir  .#cli-appimage   ->  logoscore
  .#ctl  .#ctl-bundle-dir  .#ctl-appimage   ->  logosctl
  .#     (default)                          ->  logoscore

So anything already pointing at the default or at `.#cli` -- including
every doc-test across the workspace that does
`nix build github:logos-co/logos-logoscore-cli` -- keeps getting the
tool it gets today, and logosctl is strictly opt-in.

They still compile together, since they share everything but main.cpp;
only the packaging is split. modules-pkg/ ships solely in the ctl
outputs, because logoscore never scans it.

logoscore's desktop entry and icon are restored, and logosctl gets its
own. The logosctl doc-tests now build .#ctl / .#ctl-bundle-dir.

Verified: every output builds and ships only its own binary; both
portable bundles run and report the module set expected of each.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(config): let each session subdirectory be redirected

The session directory being self-contained is what makes it portable, and
that should stay the default -- but it was also the only option, which
made reasonable setups impossible: sharing one keyring across sessions,
putting the .lgx cache on a bigger disk, or pointing at a modules tree
something else manages.

A `dirs:` block now redirects any of them:

    dirs:
      keyring: ~/.config/logos/trusted-keys
      cache: /var/cache/logos
      modules: /opt/logos/modules
      plugins: plugins-custom
      data: /var/lib/logos/data

The form of the value decides whether portability survives, which is the
part worth knowing:

    plugins-custom  -> <session>/plugins-custom   still portable
    ~/x             -> $HOME/x                    outside the session
    /var/cache/...  -> as given                   outside the session

`~` is handled because it is the natural thing to write in a config file
and would otherwise resolve to <session>/~/... , which exists nowhere.

Overrides resolve once, when set, so relocating a session afterwards
cannot silently drag an absolute path along with it. Only the daemon
applies them, and it does so before anything asks Config for a path.

persistence_path is folded into dirs.data -- it was already the same
setting under an older name -- so the two no longer need choosing
between at the point of use.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): rotating log file with configurable size and retention

--detach used to dup2 stdout/stderr straight onto daemon/daemon.log,
which grew without bound and had no rotation. A long-lived daemon needs
better than that.

There is now a logs/ directory, like Basecamp's, and a logging block:

    logging:
      enabled: true          # false -> no log file at all
      file: daemon.log       # inside dirs.logs
      max_size_mb: 10        # rotate past this; 0 = never rotate
      max_files: 5           # keep this many in total
      console: true          # mirror to the terminal

dirs.logs joins the overridable session directories, so logs can be
shipped somewhere a collector already watches.

Capture is pipe-based rather than a file redirect, and that is the
load-bearing decision: module hosts are separate processes holding
inherited descriptors. Redirecting to a file catches their output but
makes rotation impossible -- renaming a file out from under a child that
has it open just keeps filling the old inode. A pipe puts one reader in
charge, so rotation is safe and subprocess output still lands in the
log. Same shape as basecamp's LogRedirector, which solved this already.

The size cap and retention come from spdlog's rotating sink rather than
being hand-rolled; liblogos already logs through spdlog. Lines arriving
from the pipe already carry their own timestamp and level, so the sink
uses a raw pattern instead of stamping them twice.

Two bugs found while testing it:

  - Draining raced shutdown. stop() cleared the running flag before
    restoring the descriptors, so a reader holding data would process
    it, loop, see the flag clear and exit -- dropping whatever was still
    in the pipe. The last lines before a shutdown are exactly the ones
    worth keeping. EOF is now the only stop condition.
  - --detach reported the wrong path. The parent prints before the child
    has read the config, so it guessed the default and lied to anyone
    who had redirected dirs.logs or renamed the file. It now reads the
    same config the child will.

Verified live: default, disabled, and redirected-with-custom-filename
all behave and are reported accurately. Four unit tests cover capture
of both streams, no double-stamping, rotation with retention, and
disabled-is-not-an-error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(daemon): timestamp log files, and bound the directory

Adopts basecamp's naming -- each start writes its own
daemon_<yyyymmdd_HHMMSS>.log -- so a session's output is one file you can
point at, instead of every run appending into the same daemon.log.

Two things beyond copying basecamp:

  - `logging.file` survives as a symlink to whichever file is current, so
    `tail -F logs/daemon.log` follows across restarts and nobody has to
    work out a stamp. It also means --detach can report a path that is
    always valid; previously it had to guess one, and guessed wrong for
    anyone who had redirected dirs.logs.

  - max_files now bounds the *directory*, pruning oldest-first at each
    start. spdlog's retention only prunes within one sink's rotation set,
    and every start opens a new stamped base name, so without this a
    daemon restarted a hundred times would leave a hundred logs behind.
    basecamp has exactly that problem.

Verified live: three restarts leave three stamped files with the symlink
tracking the newest; five restarts with max_files: 2 leave two.

Two new tests cover the naming and the symlink resolving to the current
session, and the cross-session pruning. The rotation test needed fixing
too -- it counted the symlink as a log file, which predated the symlink
existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: ignore suffixed nix out-links

.gitignore listed `result` but not `result-*`, so every out-link from a
targeted build -- `nix build '.#ctl' -o result-ctl`, `-o result-tests`,
and so on -- was untracked-but-not-ignored, and `git add -A` committed
them as symlinks into /nix/store.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: keep releasing logoscore, and release logosctl beside it

The earlier rename left the release workflow building the `cli-*`
outputs -- which are logoscore -- while naming every artifact
`logosctl-*`. A release/** push would have shipped logoscore binaries
under the wrong name, and stopped releasing logoscore under its own.

Both are now built and published as separate, correctly-named assets:

  logoscore-{x86_64,aarch64}-linux.tar.gz   from .#cli-appimage
  logoscore-aarch64-macos.tar.gz            from .#cli-bundle-dir
  logosctl-{x86_64,aarch64}-linux.tar.gz    from .#ctl-appimage
  logosctl-aarch64-macos.tar.gz             from .#ctl-bundle-dir

logoscore's asset names are exactly what they were, which matters:
release sets fetch this repo and expect a bundle containing
`bin/logoscore`. Each tool builds from its own flake outputs, so an
asset labelled logoscore contains logoscore and nothing else.

Both jobs gained a tool matrix with fail-fast disabled, so a failure in
the under-validation logosctl cannot block a logoscore release. The
release job now collects artifacts by pattern instead of naming each
one, so retiring logoscore later means deleting a matrix entry rather
than unpicking a download list.

Release notes lead with logoscore as the tool to use, and say the two
share no state so installing logosctl cannot disturb an existing setup.

The doc-tests workflow globs doctests/*.test.yaml, which now covers both
suites, so it is no longer named after one of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: run both tools' suites, in parallel

While both binaries ship, both get tested. logoscore had no automated
coverage on this branch at all -- only its doc-tests -- so a change to
the shared runtime could regress the tool people actually use and
nothing would say so.

tests/test_cli_logoscore.cpp and tests/test_integration_logoscore.cpp
are copies of the suites frozen against logoscore's surface. Copies
rather than a parameterised shared suite on purpose: the two surfaces
genuinely differ, and this way retiring logoscore is a delete rather
than an unpick.

checks.tests-logosctl and checks.tests-logoscore are separate
derivations, so nix builds them concurrently; checks.tests aggregates
both, keeping `nix build .#checks.<sys>.tests` working for CI while now
covering both tools.

It immediately earned its keep, catching three regressions:

  - The integration harness still passed -m, which logosctl no longer
    accepts, so its daemon never started and seven integration tests
    were failing on this branch. It now writes the modules_dirs config
    the daemon reads.

  - `logoscore --version` reported "logosctl version ...". The version
    banner had been renamed wholesale; each front-end now names itself.
    Exactly the sort of thing nobody notices until a bug report cites
    the wrong tool.

  - The new log sink only mirrored to the console when stdout was a
    TTY, so `logoscore -D > logs.txt` -- which the doc-tests do --
    produced an empty file. Mirroring now follows the configured
    setting, pipe or terminal alike, and the log file is gated to
    logosctl so logoscore's output behaviour is untouched.

Both suites green: logosctl 138 unit + 18 CLI + 18 integration,
logoscore 20 CLI + 18 integration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(package): honour -o, and stop parsing command lines backwards

`package download -o DIR` accepted the flag and threw it away -- the
argument was parsed into a variable and then explicitly discarded with
`(void)outDir;`. The file went to $TMPDIR regardless. The config's
`dirs.cache` had the same problem from the other end: the directory was
created and documented as holding downloads, and nothing ever wrote to
it.

The cause was the same for both. package_downloader takes no
destination, so the file lands in $TMPDIR on the DAEMON's filesystem --
which is where the move has to happen too. Doing it client-side would
work only for a local daemon. So `downloadPackage` joins the daemon-side
package operations: it downloads, then moves the result into the
requested directory, or into the session's cache/downloads when no -o
was given. The client resolves a relative -o against its own working
directory first, so a local daemon does what the user typed; against a
remote one the path is remote, and a bad one fails loudly rather than
quietly writing elsewhere.

Writing the first test for it turned up something worse. CLI11's
`parse(std::vector<std::string>&)` consumes the vector from the BACK --
only the rvalue overload reverses for you -- so passing natural order
parses the command line backwards. `watch` and `issue-token` did reverse
first; nothing else did. It goes unnoticed with one positional and flags
(order does not matter), and is quietly wrong the moment an option takes
a value, because the option pairs with the token to its LEFT:

  package download pkg -o dir   ->  name="dir", output="pkg"
  package install a b --version 1.0
                                ->  names=["1.0","b"], version="a"

So `package install`, `search --category`, and `download -o` all
misparsed. Every site now goes through one `parseArgs` helper that
reverses, which fixes the broken ones, is a no-op for the harmless ones,
and removes the trap for the next command.

PackageCommand had no unit tests at all, which is why a discarded flag
survived review. Four now cover download; the two asserting -o reaches
the daemon fail against the old code.

142 unit + 18 CLI + 18 integration green for logosctl, 20 + 18 for
logoscore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: one README about the repo, one document per tool

The README had grown into a logosctl manual with a banner on top telling
logoscore users that everything below did not apply to them, and pointing
them at doc-test YAML for their actual documentation. Since logoscore is
still the tool to use, its documentation should not be the thing you are
told to skip.

So: README.md covers what is true of both -- what the repo is, the two
binaries and how they differ, the flake outputs, the test targets,
dependency resolution, platforms -- and hands off to one document per
tool.

  docs/logoscore.md   the usage material, unchanged, as its own document
  docs/logosctl.md    sessions, config, logs, packages, examples

Writing logosctl's own document exposed a gap: it had no command
reference at all. The rewrite dropped the client-command list, argument
typing and exit codes, and left behind a "see Argument typing below"
pointing at a section that no longer existed. All three are back, with
the command list written against the grammar that is actually
implemented (checked against normalizeGroupVerbs and the subcommand
dispatch, not from memory), plus the two defaults worth stating up front
-- install does not load, remove takes dependents.

Also fixes stale copy that survived the earlier rewrite: `load-module`
where logosctl says `module load`, and a "multiple module directories"
caption over a --config-dir example, from a flag logosctl does not have.

Deleting logoscore later is now deleting one file and a table row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(daemon): make TLS configurable again, and say why startup failed

Three bugs, all found by running the doc-tests I had just rewritten
instead of trusting them.

**tcp_ssl could not be configured at all.** `transportFromJson` never
read `cert` or `key`. That was harmless while those arrived via
`--module-transport ...,cert=...,key=...`, parsed by the CLI mini-grammar
-- but that grammar is gone, and the config file is now the only place to
set them. So every tcp_ssl listener bound with no certificate: the daemon
started, reported itself healthy, accepted connections, and failed every
handshake with "no shared cipher (SSL routines)". The client just saw
"core_service not reachable".

The stripping was deliberate but applied one layer too high: cert and key
have no business in state.json, which clients read, but the config file is
where an operator *authors* them. `transportToJson` now takes
`includeSecrets` -- true writing the config, false writing state.json. A
test asserts the round-trip, and another asserts the key path never
appears in state.json.

**`--detach` swallowed the reason startup failed.** Config validation runs
before LogSink opens the log, and the child's stderr went to /dev/null, so
a rejected config produced "daemon exited during startup. See
<path>/logs/daemon.log" -- naming a file that had never been created. The
child's early output now goes to a startup file the parent reads and
prints on failure, removed either way. LogSink takes those descriptors
over as soon as it starts, so the file only ever holds pre-logging output.

**The plaintext-TCP guard advertised a flag that does not exist.** It said
"pass --insecure-tcp"; logosctl has no such flag. It now names the config
key, `insecure_tcp: true`.

Verified end to end against a real daemon: plaintext guard refuses and
says why, loopback TCP binds and serves `status`/`module ls` from a
separate client session, TLS serves the same over 6443/6444, and dropping
the CA while keeping verify_peer still fails closed.

144 unit + 18 CLI + 18 integration green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(flake): give autoPatchelf the libraries both binaries now link

Every Linux build failed:

  auto-patchelf could not satisfy dependency libyaml-cpp.so.0.8
  wanted by .../bin/.logoscore-wrapped

The packaging derivations listed only Qt in buildInputs, which is what
autoPatchelfHook resolves DT_NEEDED entries against. yaml_json.cpp and the
log sink are in the shared sources, so *both* binaries link yaml-cpp and
spdlog -- including logoscore, which is why its Linux build broke too on a
branch that was supposed to leave it alone.

macOS does not patchelf, so this was invisible locally and in the macOS
CI jobs; only the Linux matrix caught it, and it took down the AppImage
builds, the CI job, and every Linux doc-test with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(doctests): bring the logosctl specs up to what logosctl does

Nineteen doc-test steps were failing. None of them were runtime bugs in
the specs' own right -- they were specs still describing an older
logosctl, which is its own kind of failure: a doc-test that lies is worse
than no doc-test.

  transports      Still drove `--module-transport` and hand-written
                  client/config.json. The flags had been dropped from the
                  `run:` lines but no config step replaced them, so the
                  daemon never bound TCP at all and every step after it
                  failed. Rewritten around `daemon config set` /
                  `client config set` with YAML documents, for both the
                  plaintext and TLS halves.

  daemon          Read the log at session/daemon/daemon.log; logs moved to
                  session/logs/. The crash-recovery step passed `-m`,
                  which logosctl does not accept, so its daemon never
                  started and the step reported LEAKED against a worker
                  that had never existed.

  modules-bundle  Asserted all three modules in result/modules. The
                  package modules live in modules-pkg/ so that logoscore's
                  modules/ stays byte-identical -- which the spec is now
                  the place that explains.

  packages        Expected the interactive wording ("dry run",
                  "Installed:"). Doc-tests are not a terminal, so every
                  command renders JSON. The install was working the whole
                  time; only the assertions were wrong. They now match the
                  JSON, and the prose says why it is JSON.

Rewriting the transports spec is what turned up the TLS and --detach bugs
fixed in the previous commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): a typo must not abort the daemon, and a key must not lie

Three defects in the YAML config path, all found by building a Python
client against this CLI and checking its assumptions against the binary
rather than the docs.

**A config typo aborted the process.**

  printf 'version: 2\nmodules_dirs: /single/path\n' > bad.yaml
  logosctl --config-dir ./s daemon config set ./bad.yaml
  => libc++abi: terminating due to uncaught exception ...
     [json.exception.type_error.302] type must be array, but is string

nlohmann's `json::value(key, default)` THROWS when the key is present
with the wrong type, every config read used it, and nothing caught it.
So it was not one key -- it was every key in both readers. A scalar
where a list belongs is an ordinary mistake and it killed the binary.

Now a type-checked reader (src/json_schema.h) records
"<dotted.path>: expected <what>, but got <what>" and the document is
refused whole, the same shape as the existing unknown-key error:

  {"code":"INVALID_CONFIG",
   "message":"modules_dirs: expected a list of strings, but got a string."}

Both readers went through it, including two paths that could abort the
daemon mid-boot rather than at `config set`.

**`config set` validated after writing.** A schema-invalid document was
installed and then reported as an error, leaving the session holding a
config the daemon would refuse to boot from. Validation now happens
entirely in memory first, on both the daemon and client sides -- the
client side had no schema validation at all -- and the write is
temp-file + rename instead of truncate-in-place.

That exposed a fourth: `yaml_json::dump` emitted numeric-looking strings
bare, so `port: "6001"` came back as the number 6001. The bytes
validated were not the bytes written.

**Two keys were accepted, stored, and never applied.**

`signature_policy` sat on the allowlist and was written verbatim to
config.yaml but was never even parsed. An operator setting `require` got
no enforcement and no warning. It is now parsed with a strict allowlist
and pushed into package_manager at boot beside setKeyringDirectory --
the module has had setSignaturePolicy all along. Unset issues no RPC, so
the module keeps its own default instead of having it restated.

The top-level `ssl: {cert, key, ca}` block was parsed into DaemonConfig
and read by nobody; only per-listener cert/key reached the transport
set. Configuring TLS the obvious way therefore produced listeners with
no certificate and "no shared cipher" on every handshake -- the same
failure fixed one layer down last commit. It is now a session-wide
default that per-listener values override.

Also: docs advertised `module load --no-deps`, which does not exist --
`module load` takes only a positional name and always resolves
dependencies. Corrected, along with the rest of the command reference,
verified against the binary.

logoscore is untouched: 20 CLI + 18 integration, exactly as before.
logosctl 171 unit (was 144) + 25 CLI (was 18) + 18 integration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(detach): re-exec the launcher, not the ELF it hides

`daemon start --detach` was dead on Linux portable builds. The daemon
exited immediately with status 127, no output, and no log file -- so the
only diagnostic was "daemon exited during startup. See <path>", naming a
file that had never been created.

strace, on a real Linux box, said it in one line:

  execve(".../bin/.logosctl.elf", [...]) = -1 ENOENT
  exit_group(127)

A portable bundle installs the CLI as a launcher script beside a hidden
companion:

  bin/logosctl        the launcher, a shell script
  bin/.logosctl.elf   the real ELF

The launcher exists because that ELF cannot be started on its own: its
PT_INTERP names a dynamic loader that is not on the host, so the launcher
runs it through a known-good ld.so instead. The ENOENT is the kernel
reporting the missing *interpreter* -- the ELF is right there.

--detach re-execs itself (it has to: macOS forbids running a forked
process that has initialized CoreFoundation), and it re-exec'd
executablePath(), which is that ELF.

My first attempt preferred argv[0], reasoning that it is what the caller
actually typed. That was wrong, and the trace showed it failing
identically: the launcher execs ld.so with the ELF, ld.so drops itself
from argv, and the program sees the ELF as argv[0] too. Neither source
of truth names the launcher.

So the mapping is applied to whatever candidate we end up with, using the
convention the launcher script itself documents -- the install dir is the
one holding the hidden companion `.$BASE.elf`. `bin/.logosctl.elf` maps
back to `bin/logosctl`. argv[0] is still preferred over
executablePath() (it is what was invoked, and it is right when a bare
name resolves through PATH), and it is absolutised, since the daemon may
run from a different directory.

Only this combination was ever broken: portable AND Linux AND --detach.
macOS bundles a real binary with qt.conf and no launcher, Linux dev
builds are ordinary ELFs, and the foreground -D path never re-execs. The
one doc-test that uses the portable bundle is the packages spec, and
cachix served a permanent 522 for one of its store paths from the day it
was written -- so its 14 cascading failures read as infrastructure until
the cache recovered and the real failure surfaced underneath.

Verified on Linux against the same bundle that failed: daemon starts
detached, all three bundled modules load, `daemon stop` returns ok.

Also here, and what made the diagnosis possible: --detach now prints the
TAIL of the daemon log rather than its path. The startup file only holds
output from before LogSink takes the descriptors, so a daemon that dies
after logging is up left it empty and the reason unread. That there was
no log at all is what pointed at exec.

179 unit tests (8 new, covering the launcher mapping and its edges: no
sibling, an ordinary foo.elf, a non-executable candidate, absent argv[0]).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* build: bundle logosctl/logoscore as headless Qt programs

* build: bump nix-bundle-dir and nix-bundle-appimage to main

Picks up the merged trampoline drop: per-arch psABI PT_INTERP, DT_RPATH,
qtCliApp for headless Qt, and the AppImage consumer that already tracks
the same pin. nix-bundle-dir 4fd87d1 (PR tip) → cb9afc8; appimage
8fcc56b → 04a3cf8.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-03 21:49:15 -03:00