Repository navigation
agent: kernel-locked manifest writes off the runtime, LISTEN_PID required, one frame deadline - #139
Merged
Merged
Conversation
…ired, one frame deadline Re-audit of #135: - The manifest cache's write lock was a create-new lockfile broken once older than 10 s: two waiters could both judge it stale, remove and recreate it, and both hold it. It is now the repo's session lock (`session_store::acquire_session_lock`) on `<manifest>.lock`, a kernel file lock that dies with its holder, so there is nothing to judge stale. The lock wait blocks, so `store`/`forget`/`store_skills` are async and run it with `spawn_blocking`; they were called from connection tasks on tokio workers. - `serve` adopted socket-activated descriptors with `LISTEN_PID` unset (`listenfd` only checks it when present), and the comment claimed the guard was enforced. Production now requires `LISTEN_PID` to name this process and refuses otherwise at startup. The test harness's `HeldPort` starts `serve` the way an activator does (a shell exports its own pid and execs in place), so the binary has no test-only path. - One frame deadline (`FRAME_DEADLINE`, 60 s) for every `serve` reader: `serve_frames` for stdout in all serve_* suites, `ws_next_frame` for WebSockets, and `skills_env::Serve`, which is now built on it. A stalled run fails promptly instead of at the runner's kill; `serve_harness_deadlines` holds every serve_* suite to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
…D operator note Audit of #139: - Operators: ARCHITECTURE.md's transport section says what changed for launchers that pass sockets without LISTEN_PID (`systemfd --no-pid`, forking wrappers), the error they now see, and the fix (`exec` the daemon, or set LISTEN_PID in the child just before it runs). - The frame-deadline lint matched `BufReader::new(` and `stdout` on one line. Now every child-stdout read in the serve_* and mcp_* suites and tests/common goes through `common::child_frames`, the events fixture included, and the lint fails on any other `ChildStdout` or `.stdout.take()`/`as_mut`/`as_ref` (comments dropped, whitespace ignored, so formatting cannot hide one). It found one more raw read. - `update_store_off_runtime` logs a JoinError (a panicking manifest write) at warn instead of discarding it. - The kernel lock and same-process path registry are their own module, `file_lock`; the session lock and the manifest lock both use it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
…D operator note Audit of #139: - Operators: ARCHITECTURE.md's transport section says what changed for launchers that pass sockets without LISTEN_PID (`systemfd --no-pid`, forking wrappers), the error they now see, and the fix (`exec` the daemon, or set LISTEN_PID in the child just before it runs). - The frame-deadline lint matched `BufReader::new(` and `stdout` on one line. Now every child-stdout read in the serve_* and mcp_* suites and tests/common goes through `common::child_frames`, the events fixture included, and the lint fails on any other `ChildStdout` or `.stdout.take()`/`as_mut`/`as_ref` (comments dropped, whitespace ignored, so formatting cannot hide one). It found one more raw read. - `update_store_off_runtime` logs a JoinError (a panicking manifest write) at warn instead of discarding it. - The kernel lock and same-process path registry are their own module, `file_lock`; the session lock and the manifest lock both use it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
…D operator note (#141) * agent: file_lock module, unreachable deadline-free readers, LISTEN_PID operator note Audit of #139: - Operators: ARCHITECTURE.md's transport section says what changed for launchers that pass sockets without LISTEN_PID (`systemfd --no-pid`, forking wrappers), the error they now see, and the fix (`exec` the daemon, or set LISTEN_PID in the child just before it runs). - The frame-deadline lint matched `BufReader::new(` and `stdout` on one line. Now every child-stdout read in the serve_* and mcp_* suites and tests/common goes through `common::child_frames`, the events fixture included, and the lint fails on any other `ChildStdout` or `.stdout.take()`/`as_mut`/`as_ref` (comments dropped, whitespace ignored, so formatting cannot hide one). It found one more raw read. - `update_store_off_runtime` logs a JoinError (a panicking manifest write) at warn instead of discarding it. - The kernel lock and same-process path registry are their own module, `file_lock`; the session lock and the manifest lock both use it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk * agent tests: the guard holds the stdout pipe; lint every route to it; test the inode recheck Audit of 502bce7: - The lint was bypassable (`std::mem::take(&mut child.stdout)` survived, as would `Option::take`, or destructuring a raw `Child`). `ChildGuard` now takes the pipe off the `Child` at spawn, so `child.stdout` is `None` however it is reached; `child_frames` takes it from the guard, and `raw_stdout` exists only for tests about the pipe itself. The lint is a function over source text with tests per bypass form: the `ChildStdout` type, `Option` methods on the field, assignment, `take`/`replace`/`swap(&mut ….stdout)`, `Child { .., stdout, .. }` patterns, unguarded `.spawn()`, and `raw_stdout`; it leaves an `Output`'s bytes and harness fields named `stdout` alone. The suites outside its scope (smoke, settings_store, run_cli_flags) read through `child_frames` too; run_stdout_robustness uses `raw_stdout` for its EPIPE tests. - `file_lock`'s inode recheck is tested through `try_lock`: a test seam replaces the file between open and lock, and `try_lock` must go round again and return the lock on the file now at the path. - The lint scans directory modules (`tests/mcp_tasks_env/mod.rs`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This PR fixes the three low-severity findings from the re-audit of #135. Each fix has a test that fails without it. I reverted each fix in a scripted mutation run, and all 5 reverts were caught.
session_store::acquire_session_lock, on<manifest>.lock. It is a kernel file lock that dies with its holder, so a crashed writer's leftover file is simply unlocked and there is no staleness to judge.mcp_manifest::tests::a_dead_writers_leftover_lock_admits_exactly_one_holder_at_a_time: 16 waiters race for a lock file with an old timestamp, 20 rounds, and at most one may hold it at a time. Restoring the old age-based break fails it.update_storewas called from connection tasks on tokio workers.store,forgetandstore_skillsare nowasyncand run the locked read-modify-write withspawn_blocking. Callers inmcp.rsandmcp_skills.rsawait them.mcp_manifest::tests::a_write_waiting_for_the_lock_does_not_block_the_runtime: on a single-threaded runtime, a write that waits 400ms for the lock still lets a 10ms ticker run. Calling the write directly on the runtime fails it.LISTEN_PID == getpid()guard was enforced, butlistenfdskips the check whenLISTEN_PIDis unset.serve_ws::check_listen_pidrefuses at startup, with a clear error, whenLISTEN_PIDis unset or names another process. Every real activator (systemd,systemd-socket-activate) sets it, so an unset one means the variables did not come from activation. The docs now match. The test harness'sHeldPortkeeps working through a test-only mechanism outside the binary: it startsservethe way an activator does, throughsh -c 'LISTEN_PID=$$; export LISTEN_PID; exec "$0" "$@"'(execkeeps the pid), so the binary has no test-only path.serve_systemd::serve_refuses_passed_sockets_without_its_own_listen_pid(unset, and a foreign pid:serveexits namingLISTEN_PID; removing the check makes it adopt and serve).serve_systemd::serve_adopts_a_passed_socket_whose_listen_pid_is_its_ownand the unit testpassed_sockets_are_taken_only_when_listen_pid_names_this_process. The existingserve_adopts_a_systemd_activated_socket(realsystemd-socket-activate, present on this host) still passes, as do all the MCP Events suites that useHeldPort.serve_state_reporting.common::FRAME_DEADLINE(60s) is used by everyservereader:common::serve_framesfor stdout (240 readers across the 25serve_*suites that read stdout),common::ws_next_framefor WebSockets (so every WS suite is covered), andskills_env::Serve, which is now built on the sameFramesreader instead of its own copy of the logic. A stalled run fails with a message saying why, instead of waiting for nextest's kill.serve_harness_deadlines: a silent stdout and a silent WebSocket each fail at their deadline (removing the WS timeout makes the read wait 30s), andevery_serve_suite_reads_frames_with_the_shared_deadlinefails if anyserve_*suite orskills_envreads stdout through a bareBufReader.Proof (local, 6a9898f)
beyond-ai-agenttest exceptexec_endpoint_live(lib, allmcp_*, allserve_*): 2512/2512 passed.-D warnings, agent + test-support + fleet-sim, all targets, code-mode),cargo fmt --checkanddprint check: clean.CI: green on the first attempt: every job passed (run 37574497085), including agent-serve, agent-rest, agent-lib, agent-code-mode and the changed-lines mutants.
Notes
<manifest>.locknow stays on disk between writes (a kernel lock needs a file to lock; removing it would reopen the race this fixes). It is empty.servelaunched withLISTEN_FDSbut without its ownLISTEN_PIDused to adopt whatever sat at fd 3. It now refuses to start. Only an environment that was never real socket activation is affected.🤖 Generated with Claude Code
https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk