Runs outlive the agent: run supervisor, run directories, adoption, late effectful verdicts (SPEC §7.5) #28

Manually merged
krisbuild merged 8 commits from lu/run-supervisor into main 2026-09-27 23:28:17 +02:00
Owner

Builds on #21 (acked delivery), which builds on #12 and #15 — until those land, this diff includes them.

What

Live-upgrades §4.1 ("runs outlive the agent") and the late-effectful-verdict part of §4.3, specified as the new SPEC §7.5 with amendments to §3.3, §3.3.1, §6.1, §6.2, §7.1, §7.2, §7.2.1.

  • Run directory <work_dir>/runs/<instance>-<attempt>/: run.json (contract version, assignment, run token, lease inputs, launch, supervisor pid + start time, slot, phase, acknowledged log offset, verdict), supervisor.json, task.json, a framed log ([stream u8][len u32 LE][bytes], whole UTF-8 per frame, secrets masked), exit.json (renamed in only once the tree is dead), a cancel request file (class inside), out/stdout, w/, secrets/. One writer per file; fields only grow, anything else bumps CONTRACT.
  • kb-agent supervise <run-dir> (same binary, own session): spawns the task and sidecars, is its subreaper and owns its cgroup (teardown moved here from proc.rs, which is gone), holds the slot flock by inherited fd, writes the log, honours an absolute deadline and cancel requests, writes exit.json only after teardown. No network, no protocol. exec: direct child; container: podman under the supervisor (rm -f on teardown, waited for); nix-drv: the nix build client; microVM: cloud-hypervisor + two virtiofsd sidecars, the guest's exit file ends the run. In-process agents (kb run, tests) run the same supervisor code as a tokio task.
  • The agent drives: tails the log into Log events and persists the offset its acknowledged events cover; renders the NixDrv realiser log (#25) and folds daemon builds into usage; relays cancel/timeout/lease-loss as cancel requests; after exit.json runs the post-phases (collected → pushed → reported), each recorded and idempotent; sends Finished (or the negotiated legacy sequence) and removes the directory once that event is retired.
  • Adoption before the first Hello: supervisor alive (pid + start time) → drive on from the acked log offset; exit.json present or phase past running → resume post-phases; neither → forget, killing anything a crashed supervisor left. Hello.leases claims adopted and unacked runs with additive leases_complete: true; a CP seeing it takes back at once what the Hello leaves out (assigned in an earlier second): re-queued, or lost for effectful — which lifts the rerun refusal.
  • systemd: agent unit Delegate=yes, DelegateSubgroup=agent, KillMode=process (module checks); runs in <unit>/runs/<token>; old fallback roots without delegation.
  • Late effectful verdicts (CP): a Finished for an effectful instance failed lease-expired, from its assigned node, replaces the failure (through #8's infra-aware finish, logged). A late success returns dependency-failed dependents to pending and reopens a failed graph so it completes and posts anew; cancelled/superseded graphs are left alone. The rerun refusal stays while the run may still be going; the instance stays charged to its node until its timeout.
  • Failure classes (#8 table): spawn-error only when nothing ran; wait-error a failed wait; new supervisor (supervisor died with the task started — infra, against the node) and lost (infra, not against the node).

Why

Every agent restart (any nixos-rebuild switch touching the package) killed every task on the node and failed effectful ones outright. After this, a restart costs a run a few seconds of log latency.

Deploy notes

  • The first switch to this build still restarts the old agent the old way (the old unit has no delegation). Deploy while idle or drain first. Every later switch is harmless.
  • Upgrade order unchanged (CP first). No new protocol generation; an older CP ignores leases_complete and reaps forgotten runs at lease expiry as before.
  • To stop a node for good: drain, then stop.
  • Operator guide: "Agent restarts (SPEC §7.5)" in docs/forgejo-setup.md.
  • Only a real systemd deploy proves the delegation path (DelegateSubgroup needs systemd ≥ 254; nixpkgs ships 257), cgroup +cpu +memory enablement under User=krisbuild, and survival across systemctl restart; tests and CI exercise the non-delegated fallback plus in-process and real-subprocess adoption.

Tests

Supervisor (exit.json only after tree death incl. orphans, with/without cgroup; cancel/deadline teardown; masking; spawn failure); run-dir contract round-trip and pinned contract-1 shape; log tail/ack/resume; in-process adoption (agent dropped mid-run → attempt 1, contiguous log; exit.json present; reboot leftovers forgotten); subprocess adoption with real kb-agent supervise (effectful run survives its agent; adopted run can be cancelled); CP it/adoption.rs (complete Hello re-queues the unclaimed at once; late verdict replaces lease-expired, dependent revived, graph failed → succeeded, deploy never ran twice; complete Hello lifts rerun refusal); health classes; module check.

🤖 Generated with Claude Code

**Builds on #21 (acked delivery), which builds on #12 and #15** — until those land, this diff includes them. ## What Live-upgrades §4.1 ("runs outlive the agent") and the late-effectful-verdict part of §4.3, specified as the new **SPEC §7.5** with amendments to §3.3, §3.3.1, §6.1, §6.2, §7.1, §7.2, §7.2.1. - **Run directory** `<work_dir>/runs/<instance>-<attempt>/`: `run.json` (contract version, assignment, run token, lease inputs, launch, supervisor pid + start time, slot, phase, acknowledged log offset, verdict), `supervisor.json`, `task.json`, a framed `log` (`[stream u8][len u32 LE][bytes]`, whole UTF-8 per frame, secrets masked), `exit.json` (renamed in only once the tree is dead), a `cancel` request file (class inside), `out/stdout`, `w/`, `secrets/`. One writer per file; fields only grow, anything else bumps `CONTRACT`. - **`kb-agent supervise <run-dir>`** (same binary, own session): spawns the task and sidecars, is its subreaper and owns its cgroup (teardown moved here from `proc.rs`, which is gone), holds the slot flock by inherited fd, writes the log, honours an absolute deadline and cancel requests, writes `exit.json` only after teardown. No network, no protocol. exec: direct child; container: podman under the supervisor (`rm -f` on teardown, waited for); nix-drv: the `nix build` client; microVM: cloud-hypervisor + two virtiofsd sidecars, the guest's exit file ends the run. In-process agents (`kb run`, tests) run the same supervisor code as a tokio task. - **The agent drives**: tails the log into `Log` events and persists the offset its acknowledged events cover; renders the NixDrv realiser log (#25) and folds daemon builds into usage; relays cancel/timeout/lease-loss as `cancel` requests; after `exit.json` runs the post-phases (`collected` → `pushed` → `reported`), each recorded and idempotent; sends `Finished` (or the negotiated legacy sequence) and removes the directory once that event is retired. - **Adoption before the first Hello**: supervisor alive (pid + start time) → drive on from the acked log offset; `exit.json` present or phase past running → resume post-phases; neither → forget, killing anything a crashed supervisor left. `Hello.leases` claims adopted and unacked runs with additive `leases_complete: true`; a CP seeing it takes back at once what the Hello leaves out (assigned in an earlier second): re-queued, or `lost` for effectful — which lifts the rerun refusal. - **systemd**: agent unit `Delegate=yes`, `DelegateSubgroup=agent`, `KillMode=process` (module checks); runs in `<unit>/runs/<token>`; old fallback roots without delegation. - **Late effectful verdicts (CP)**: a `Finished` for an effectful instance failed `lease-expired`, from its assigned node, replaces the failure (through #8's infra-aware finish, logged). A late success returns `dependency-failed` dependents to `pending` and reopens a `failed` graph so it completes and posts anew; cancelled/superseded graphs are left alone. The rerun refusal stays while the run may still be going; the instance stays charged to its node until its timeout. - **Failure classes (#8 table)**: `spawn-error` only when nothing ran; `wait-error` a failed wait; new `supervisor` (supervisor died with the task started — infra, against the node) and `lost` (infra, not against the node). ## Why Every agent restart (any `nixos-rebuild switch` touching the package) killed every task on the node and failed effectful ones outright. After this, a restart costs a run a few seconds of log latency. ## Deploy notes - **The first switch to this build still restarts the old agent the old way** (the old unit has no delegation). Deploy while idle or drain first. Every later switch is harmless. - Upgrade order unchanged (CP first). No new protocol generation; an older CP ignores `leases_complete` and reaps forgotten runs at lease expiry as before. - To stop a node for good: drain, then stop. - Operator guide: "Agent restarts (SPEC §7.5)" in `docs/forgejo-setup.md`. - Only a real systemd deploy proves the delegation path (`DelegateSubgroup` needs systemd ≥ 254; nixpkgs ships 257), cgroup `+cpu +memory` enablement under `User=krisbuild`, and survival across `systemctl restart`; tests and CI exercise the non-delegated fallback plus in-process and real-subprocess adoption. ## Tests Supervisor (exit.json only after tree death incl. orphans, with/without cgroup; cancel/deadline teardown; masking; spawn failure); run-dir contract round-trip and pinned contract-1 shape; log tail/ack/resume; in-process adoption (agent dropped mid-run → attempt 1, contiguous log; exit.json present; reboot leftovers forgotten); subprocess adoption with real `kb-agent supervise` (effectful run survives its agent; adopted run can be cancelled); CP `it/adoption.rs` (complete Hello re-queues the unclaimed at once; late verdict replaces lease-expired, dependent revived, graph failed → succeeded, deploy never ran twice; complete Hello lifts rerun refusal); health classes; module check. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
kris added 38 commits 2026-09-27 22:35:38 +02:00
protocol: handshake, capabilities as placement, eval frontend generation (SPEC §7.1)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy failed on ares (exit-code)
krisbuild/kris/krisbuild/nix/build superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 2fed6fad6f23e67827004279885a515acee25798
395e035abb
Version skew becomes explicit and safe for the mixed CP/agent versions a
staggered nuxbox/ares upgrade produces (docs/live-upgrades.md §4.4, and the
`Hello.leases` part of §4.1).

- `Hello` carries `proto` (absent = 0, legacy), `build` (`kb_core::BUILD`:
  `<version>+<rev>` from the flake, else the crate version), `caps`, and
  `leases`, the instances the agent is running. All additive: today's
  control plane ignores them.
- The control plane answers every admitted Hello with `Welcome { proto,
  build }` (not sent to legacy agents, which could only log it as
  undecodable) and an immediate `LeaseAck` that renews the claimed leases
  exactly as a heartbeat does. It admits proto cp-1..=cp plus legacy 0 and
  refuses others with a 1008 close naming what to upgrade; refusals are
  logged and shown on /ui/nodes, /api/nodes and `kb nodes`, which also
  gain the agent build, proto and caps.
- `kb_core::caps::required_caps(def, sched)`: an exhaustive map from every
  payload/input/isolation/effect/resolved-ref kind and SchedMeta enum value
  to a cap, empty for everything at proto 1, plus `constraints.caps`. The
  matchmaker requires them (legacy agents have exactly the frozen legacy
  set) and checks resolved refs against the chosen node. The skew rule is
  written into kb-core's docs and SPEC.
- The bootstrap eval script is `KB_EVAL_FRONTEND=1 exec kb-eval-nix ...`:
  the frontend generation is identity (every eval's def_hash changes once;
  no kbN- bump), the eval requires `eval-nix:1`, agents with the nix
  feature advertise it, and kb-eval-nix (and `kb` as kb-eval-nix) refuses
  another generation. An environment variable rather than a flag, because
  today's frontend ignores it where it would reject an unknown flag - so a
  legacy agent is safely treated as having `eval-nix:1`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Control-plane restarts cost running tasks nothing (SPEC §7.1)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by e924dd9245dd0ff9eb87276b57496061dd552427
b09d68e51e
A restart (nixos-rebuild switch, crash) or a control plane down for minutes
(a bad deploy being rolled back) no longer kills or re-runs fleet work:

- Socket activation: the CP takes listen_addr/agent_listen_addr from
  systemd (LISTEN_FDS, via listenfd), matched by address, else binds. The
  NixOS module adds krisbuild-control-plane.socket when `settings` gives
  the addresses; the service requires it, restarts in place
  (stopIfChanged = false keeps the socket open) with TimeoutStopSec = 30.
- Graceful stop on SIGTERM/SIGINT: stop accepting, bounded (5 s) finish of
  in-flight requests, the scheduler pass and merge-queue tick, event
  streams end, agent sockets close with 1012, the db actor drains.
- Boot lease grant: before the first pass, lease_until = max(lease_until,
  now + reconnect_grace_secs) (default 120) for assigned/running instances.
- Agent: redials at once on 1012, else 1-5 s jittered backoff; first
  heartbeat right after Hello; a CP silent for 20 s counts as gone. While
  disconnected, unconfirmed runs survive disconnected_grace_secs (default
  900) instead of lease_ttl; the first LeaseAck of a new connection is
  authoritative.
- kb follow resubscribes to a dropped event stream (backoff, up to 15 min)
  without replaying logs already shown.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
build id at runtime via a package wrapper; keep a reconnect's registration
Some checks reported errors
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by a58cdf2619731327be18a42d7b91e392fe1333aa
2fed6fad6f
`kb_core::build()` reads `KB_BUILD` at runtime (falling back to the crate
version) instead of compiling the rev in, so `.#ci.build`, clippy and test
no longer change with every commit and a local pre-warm matches CI. The
flake's packages (and so the NixOS module's default) are now a
symlinkJoin + wrapProgram around the unchanged cargo build that sets
`KB_BUILD=<version>+<rev>` by default; wrapProgram inherits argv0, so the
PATH-resolved `kb-eval-nix` and `kb` keep dispatching on their names.

A connection's disconnect cleanup now removes the node's registry entry
(and marks it offline) only while that entry is still its own, so an agent
that reconnected before its old socket's close was processed stays
registered. docs/TODO.md tracks dropping the legacy proto-0 admission.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A legacy (proto 0) Hello renews nothing, so it no longer gets the
  immediate LeaseAck: that ack would stamp every held instance as confirmed
  while its lease ran out, letting the agent outlive a re-queue by up to a
  lease_ttl. Its first heartbeat renews and is acked as before; the
  handshake test now checks a reconnecting legacy agent gets no reply.
- Inputs are resolved again only once a node's semaphores are held (once
  per candidate); the resolved-ref cap check stays before dispatch and
  gives the holds back if the node cannot decode them.
- required_caps is computed once per candidate and passed to admissible /
  admit_reason instead of per (task, node).
- `ci.deployed` puts the NixOS module's wrapper package under CI's eval;
  corrected the argv0 comment (wrapProgram inherits it).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main into lu/protocol-handshake
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy failed on ares (exit-code)
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue head graph failed
a58cdf2619
CP restarts: semaphore/effectful lease rules, two-stage stop, review fixes
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 63f0e903ca53f98f5b822c09cc8e0ec3f2a03a19
e924dd9245
- Agent lease rules per run: semaphore holders abort at lease_ttl even while
  disconnected (the CP hands their holds on at reap, SPEC §6.4); effectful
  runs without semaphores are never aborted for lease loss (never
  re-queued, SPEC §3.3); everything else keeps the disconnected grace.
- Watchdog reads its allowance before the confirmation stamp.
- Stop in two stages: listeners, scheduler, merge queue and event streams
  first; then 1012 to agents, applying their messages until the close echo.
- A failed listener exits non-zero; LISTEN_FDS is taken before the runtime.
- Agent: silence check has its own arm, websocket writes time out after
  10 s, and 1012 redials at once only once per healthy session.
- kb follow: an outage ends only when a resubscription delivers an event;
  4xx on resubscribe fails at once.
- Docs: socket-unit restarts after an address change, stopping the socket,
  the new lease rules; tests poll instead of tight sleeps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Strict lease allowance for semaphore holders; hold reruns of lost effectful tasks
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy failed on ares (exit-code)
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test failed on ares (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue the pull-request head changed
63f0e903ca
- A semaphore holder's watchdog allows lease_ttl·3/4 − 1 s, so with its
  lease_ttl/4 tick it fires lease_ttl − 1 s after the last ack arrived:
  before the control plane's lease expires and hands the mutex on. A CP
  outage longer than that still kills such runs (accepted).
- An effectful instance failed as lease-expired may still be running on its
  node; rerunning its graph (API, UI buttons, merge-queue commands) is
  refused with 409 and a named reason until its timeout, counted from its
  start, has passed. New DbError::Conflict; UI buttons get a refusal page.
- SPEC §3.3/§6.4/§7.1 and the operator guide say so; late verdict
  acceptance is noted as the follow-up.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main into lu/protocol-handshake
Some checks are pending
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/queue queued (#15 in line)
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
63bf38a6e6
Conflict only in ws.rs imports: kept this branch's handshake imports and
main's LogStream/Payload/TaskDef.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main; virtual nodes report no host load; clippy
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/queue queued (#15 in line)
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test failed on ares (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
512a2e1beb
Merge main (verdict integrity, PR #9): SPEC §7.1 keeps both the invalid-
outputs paragraph and the control-plane downtime section.

The example runner's virtual nodes share one host, but each heartbeat
reported the host's /proc memory and load as the node's own `observed` and
`load1`, and admission charges max(reserved, observed) against the node's
declared totals. With the first heartbeat now sent right after Hello, that
reading lands before the first dispatch; on a busy CI host (memory in use
above a virtual node's 16 GB) both `warm` and `cold` looked full and
cache-affinity's e2e never left the queue. Agents now report node-wide
measurements only when the node is the host (`observe_node`; the default
`kb run` fleet), never for a virtual one.

Also: graph::effectful_still_running reads rows into a struct
(clippy::type_complexity).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	crates/kb-agent/src/conn.rs
#	crates/kb-control-plane/src/state.rs
#	flake.nix
tests: handshake suite joins the single integration binary (tests/it)
Some checks failed
krisbuild/queue the pull-request head changed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
49a4d9dc1d
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main: single integration-test binary per crate
Some checks failed
krisbuild/queue the pull-request head changed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
02484db281
tests/restart.rs moves into kb-control-plane's tests/it/ as a module; the
restart test in kb-cli's remote.rs and the harness helper followed their
files' renames. flake.nix keeps both main's agent ExecStart binding and the
socket-unit checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Acknowledged delivery: Finished, seq/acked_seq, cancel resend, DB backups (proto 2)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/deployed superseded
krisbuild/kris/krisbuild/nix/clippy superseded
krisbuild/kris/krisbuild/nix/test-deps superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/build superseded
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild/nix/workspace-deps superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 5c1e329405fbe2dcc368366b06b89dd89274e136
ddf7e444bb
Agent events carry a per-process (boot, seq); the control plane applies each
seq once and answers with LeaseAck.acked_seq, and the agent keeps events until
acked, re-sending them in order after a reconnect. A run ends in one Finished
event applied in one transaction (a success short of its outputs fails
outputs-invalid / expansion-invalid). Negotiated through Welcome: an agent
falls back to Usage/Outputs/State and delete-on-write with a proto < 2 or
legacy control plane, which in turn admits proto 0/1 agents. Cancels an agent
missed while disconnected are re-sent when it claims the instance. The control
plane copies its database to backups/ when the build changes (keeps 3), and
SPEC §9 states the expand/contract migration rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Acked delivery: a failed apply is not acked; a failed backup does not block start
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test-deps failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/test dependency-propagated failure
krisbuild/kris/krisbuild/nix/workspace-deps failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/build dependency-propagated failure
krisbuild/kris/krisbuild/nix/clippy dependency-propagated failure
krisbuild/kris/krisbuild/nix/deployed dependency-propagated failure
krisbuild/kris/krisbuild/nix/kb-check dependency-propagated failure
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
5c1e329405
An event that hits a database error while being applied no longer counts as
applied: its seq is released and the connection is closed, so the agent
re-sends it and everything after it on the next connection instead of a later
seq acknowledging it away. A database backup that cannot be written is logged
as an error and startup continues (migrations only add; SPEC §9).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- An event delivered on write (older or slow control plane) carries no seq:
  the legacy sequence splits one queued Finished into several messages,
  which a deduping control plane took for re-sends of the first. A welcome
  arriving after the settle wait switches the connection to acked delivery.
- The control plane acknowledges only a mark below every seq still being
  applied or given back, so a reconnect overlapping an apply on the old
  connection cannot ack an event that later fails.
- A failed apply closes the connection (re-send) only when a retry may help:
  busy, locked, full, I/O, actor gone. Otherwise the instance fails
  event-unappliable with the error in its log and the event is acked.
- A failed backup copy is removed so pruning does not count it; file errors
  are DbError::Io.
- Docs: SPEC §7.1 acknowledged delivery restated; rolling back across a
  protocol generation means rolling agents back too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/main' into lu/acked-delivery
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/build dependency-propagated failure
krisbuild/kris/krisbuild/nix/clippy dependency-propagated failure
krisbuild/kris/krisbuild/nix/workspace-deps failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/kb-check dependency-propagated failure
krisbuild/kris/krisbuild/nix/deployed dependency-propagated failure
krisbuild/kris/krisbuild/nix/test cancelled
krisbuild/kris/krisbuild/nix/test-deps cancelled
krisbuild/kris/krisbuild krisbuild kris/krisbuild: graph cancelled
2fefeeb755
# Conflicts:
#	crates/kb-control-plane/src/state.rs
Merge origin/main (#20)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/clippy superseded
krisbuild/kris/krisbuild/nix/build superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 733641204219f7619ff9a61bf39091e08d21d390
krisbuild/queue the pull-request head changed
30216c1968
state.rs: AppState keeps main's drvs_present and watch import alongside the
stop tokens and agent task tracker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/main' into lu/protocol-handshake
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue the merge conflicts in SPEC.md, crates/kb-agent/src/conn.rs, crates/kb-control-plane/src/http.rs, crates/kb-control-plan
7782ef75de
- When a failed apply gives a seq back on a connection the agent has already
  replaced, the newer connection (which dropped the re-sent seq as claimed)
  is closed as well, so the agent rewinds and re-sends it.
- SQLITE_PROTOCOL and SQLITE_INTERRUPT count as transient.
- A backup whose copy succeeded but whose pruning failed only warns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main (#10 enrollment, #8 node health, #22)
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test dependency-propagated failure
krisbuild/kris/krisbuild/nix/test-deps failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check failed on nuxbox (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue head failed: nix/kb-check (exit-code on nuxbox), nix/test-deps (exit-code on nuxbox)
7336412042
- kb-agent conn.rs: the reconnect loop keeps the 1012 immediate redial,
  jitter and 5 s cap, and handles an enrollment refusal first with its own
  jittered backoff capped at 30 s, as before, so a node awaiting approval
  does not poll every 5 s. session() keeps both the enrollment header and
  refusal and the silence/send-timeout handling.
- flake.nix: the nixos-module check keeps the enrolling-node eval and the
  socket-unit assertions.
- kb-cli remote tests: both the restart and the enrollment test; the
  restart test passes spawn_agent's new enroll flag.
- Lease reaping and #8's infra requeue agree: reaping still never requeues
  effectful work, and settle_failure requeues effectful only before the
  payload started, so an Exempt run is never duplicated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge origin/main into lu/protocol-handshake
Some checks reported errors
krisbuild/queue head failed: nix/test (exit-code on nuxbox)
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/clippy cached
krisbuild/kris/krisbuild/nix/build cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/kb-check cached
krisbuild/kris/krisbuild/nix/deployed cached
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by d09be7874de92b3b596b8960fc8335c587fae32d
975e0f4d67
Brings in agent enrollment (#10), node health / quarantine (#8) and the
NixOS credential fallbacks (#22). Conflicts resolved keeping both sides:
- matchmaking: required caps stay a hard constraint inside `eligible`, so
  infra-failure avoidance only counts nodes whose agent has the caps;
  admissible/admit_reason take both `caps` and `blocked`.
- load_nodes takes both the refusals and the config (health), and the
  nodes page shows refusal, health, agent build and caps rows.
- conn.rs / FakeAgent: enrollment headers and the proto-1 Hello with leases
  (dial_with_headers underneath both connect and dial).
- SPEC §6.3: caps and the health gate in the candidate predicate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/main' into lu/acked-delivery
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/deployed superseded
krisbuild/kris/krisbuild/nix/workspace-deps superseded
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild/nix/clippy superseded
krisbuild/kris/krisbuild/nix/build superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/test-deps superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by d7bebefe886c60e1d7116d375af4080d28746ce2
9bb31a6cbf
# Conflicts:
#	SPEC.md
#	crates/kb-agent/src/conn.rs
#	crates/kb-cli/tests/it/remote.rs
#	crates/kb-control-plane/src/db/mod.rs
#	crates/kb-control-plane/src/http.rs
#	crates/kb-control-plane/src/scheduler/matchmaking.rs
#	crates/kb-control-plane/src/ui/mod.rs
#	crates/kb-control-plane/src/ws.rs
#	crates/kb-control-plane/tests/it/common/mod.rs
#	crates/kb-core/src/protocol.rs
#	flake.nix
# Conflicts:
#	crates/kb-agent/src/conn.rs
#	crates/kb-control-plane/src/scheduler/matchmaking.rs
#	crates/kb-control-plane/src/ui/mod.rs
#	crates/kb-control-plane/tests/it/common/mod.rs
# Conflicts:
#	crates/kb-agent/src/conn.rs
exec: read NixDrv usage off the Finished event
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/test-deps superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by ed076c80053d40291c9ff51e5ab3ac92eea3a18b
d7bebefe88
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Every run gets <work_dir>/runs/<instance>-<attempt>/ (SPEC §7.5): run.json
  (the agent's record: contract, assignment, token, lease inputs, launch,
  supervisor, slot, phase, acknowledged log offset, verdict), a framed log,
  exit.json written once the tree is dead, a cancel request file, out/.
- `kb-agent supervise <run-dir>` starts the task in a session of its own,
  becomes its subreaper, owns its cgroup, holds the slot lock by inherited
  fd, writes the log (whole UTF-8 per frame, secrets masked), honours the
  deadline and cancel requests, and writes exit.json only after teardown.
  proc.rs goes; its teardown lives in the supervisor. Exec, container
  (podman under the supervisor, rm -f on teardown), nix-drv (the realise
  client under the supervisor) and microVM (cloud-hypervisor + virtiofsd
  sidecars, done by exit file) all launch through it. In-process agents
  (kb run, tests) run the same supervisor as a task.
- The agent drives: tails the log into Log events, persists the offset its
  acknowledged events cover, and after exit.json runs the post-phases
  (collect/upload, push), recording each; it sends Finished and removes the
  directory once that event is retired. spawn-error stays "nothing ran";
  a wait failure is wait-error, a supervisor lost without a verdict
  "supervisor".
- Adoption before the first Hello: live supervisor -> drive on from the
  acknowledged offset (slot reserved until outputs are collected);
  exit.json -> resume the post-phases; neither -> forget. Hello claims
  adopted and unacknowledged runs with the additive leases_complete; the
  control plane takes back at once what such a Hello leaves out (re-queue,
  or effectful failed lost; lease-expired effectful -> lost lifts the rerun
  refusal).
- Control plane: a late Finished for an effectful lease-expired instance
  replaces the failure; a late success revives dependency-failed
  dependents and reopens a failed graph. Such an instance stays charged to
  its node's reservation until its timeout.
- NixOS module: agent unit Delegate=yes, DelegateSubgroup=agent,
  KillMode=process; the agent keeps to <unit>/agent and puts runs in
  <unit>/runs/<token>. Module checks for them.
- SPEC §7.5 new; §3.3, §6.1, §6.2, §7.1, §7.2, §7.2.1 amended; operator
  guide, live-upgrades rollout and TODO updated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/lu/acked-delivery' into lu/run-supervisor
All checks were successful
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
ce7271a077
Brings main's node health (#8), agent enrollment (#10) and NixDrv usage
from the realiser log (#25) onto the run supervisor:
- proc.rs stays deleted; the realiser's internal-json log is rendered and
  its daemon builds accounted by the agent as it tails the run log, and
  folded into the sampler's curve once the tree is dead.
- The infra classification holds: spawn-error only when nothing ran (the
  supervisor could not start, or its exit.json says the task never did),
  wait-error for a failed wait, drv-fetch before the supervisor; the new
  `supervisor` (died with the task started) and `lost` classes are infra
  after the payload started.
- The late-verdict path goes through the infra-aware finish.
- graph.rs test: a graph row for the instances it inserts, now that the
  test schema enforces foreign keys.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
kris added 10 commits 2026-09-27 23:12:10 +02:00
Merge remote-tracking branch 'origin/main' into lu/cp-restart
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test failed on nuxbox (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue the pull-request head changed
58ac9b5bfe
graph tests: give the lease-expired rerun test its pipeline and graph rows
All checks were successful
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
ed076c8005
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The rerun refusal for a possibly running effectful task keeps working
through the forge outbox's rerun signature (the default timeout now comes
from the config it takes); a graceful stop aborts the outbox worker and
reconciliation, both durable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A late effectful verdict is recorded on its instance always, never
  requeued, and its per-task status posted (by the pass while its graph
  runs; directly for a finished graph unless a rerun reports that commit
  and context). Dependents are revived and an expansion applied only while
  the graph is still running, not superseded or cancelled, not rerun and
  with no newer graph on its (pipeline, ref). A finished graph is never
  reopened.
- Matchmaking dispatches through the connections it read before
  assigning, so an assignment stamped before a node's Hello went to the
  connection before it, which that Hello may then judge.
- An agent restarted while setting a run up reports it not-started
  (infra, before the payload: requeued, effectful included) instead of
  leaving the control plane to find it gone.
- Charging and the lost-marking scan filter lease-expired in SQL, look
  back LATE_RUN_WINDOW_SECS, and a run without any timeout is charged for
  that window, not forever.
- A damaged run record reports class lost, not supervisor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- kb-agent conn.rs: #15's Hello{leases}/Welcome and close-reason refusals
  alongside the 1012 restart, backoff, silence and enrollment handling;
  closed() maps 1012 to Ended::Restarting, a reasoned non-normal close to an
  error, anything else to Ended::Closed.
- Graceful stop covers #14's workers: reconciliation stops with the first
  stage; the outbox makes one final drain in the second, after the last
  scheduler pass and merge-queue tick, and keeps the rest queued.
- graph::rerun takes the config (#14) for the outbox write; the lease-expired
  effectful guard refuses before anything is written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
graph tests: give the lease-expired rerun test its pipeline and graph rows
All checks were successful
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue merged
14bdcfcccb
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/lu/cp-restart' into lu/acked-delivery
All checks were successful
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue merged
7d5f182d36
# Conflicts:
#	crates/kb-agent/src/conn.rs
#	crates/kb-control-plane/src/state.rs
- task.json records the launch's sidecars and teardown commands (written
  before the task starts and again after); an agent tearing down after a
  dead supervisor kills the task's group, even with its leader gone (its
  members bearing the run token), the sidecars' groups, the cgroup, and
  runs the teardown commands (podman rm -f).
- A root agent takes the delegated layout only where systemd marked the
  cgroup delegated (trusted.delegate / user.delegate), never by being able
  to write it.
- ProcRef carries the kernel boot; an in-process supervisor that fails
  writes a SupervisorLost exit so its run does not wait forever; the
  supervisor logs to supervisor.log; start timeout 60 s.
- The acknowledged log offset keeps being recorded through the
  post-phases; a resumed NixDrv past its tree renders its realiser log.
- A re-delivered assignment of a run whose verdict is pending is ignored.
- Reservation charging: a column the row reader no longer took was still
  selected, failing every pass that had such an instance.
- SPEC §7.5 / operator guide: not-started, leftover teardown, drain before
  rolling back to a build without supervisors; §6.1 late-verdict rules.
- Tests: kill_leftovers (group, sidecar, teardown command), newer contract,
  slot kept across a restart, a killed subprocess supervisor, late verdicts
  in a running and a finished graph, forget_unclaimed's cut-offs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge remote-tracking branch 'origin/lu/acked-delivery' into lu/run-supervisor
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/deployed superseded
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 7f8d388e2afc0485570609a5c848d95c9d7c61d1
bb893cebad
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Take back undelivered assignments; durable run records; late-verdict and window follow-ups
All checks were successful
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue merged
7f8d388e2a
- An assignment whose connection is gone or was replaced mid-pass is not
  sent but returned to the queue at once (holds released), instead of left
  to a lease reaper that would fail an effectful task that never ran.
- write_json fsyncs the directory after the rename, and a new run's
  directory entry is synced, so a power cut cannot bring a supervised run
  back as preparing (and so not-started and started again).
- A late verdict must be terminal; a late failure without a class is lost.
  Its per-task status is not posted when any newer graph of the same commit
  exists.
- The rerun refusal is capped at LATE_RUN_WINDOW_SECS like charging and the
  Hello's reclassification; SPEC and the operator guide say so.
- Forgetting a run skips killing leftovers recorded in another boot, and runs
  teardown commands in the background, concurrently; an in-process
  supervisor that fails has its leftovers killed before its exit is written.
- leases_complete doc wording.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner
@krisbuild r+
krisbuild manually merged commit ed82c3885c into main 2026-09-27 23:28:17 +02:00
Collaborator

Merged as ed82c3885c.

Merged as ed82c3885c73.
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kris/krisbuild!28
No description provided.