Control-plane restarts cost running tasks nothing (SPEC §7.1) #12

Manually merged
krisbuild merged 11 commits from lu/cp-restart into main 2026-09-27 23:27:54 +02:00
Owner

What

Restarting the control plane (a nixos-rebuild switch, a crash), or having it down for minutes while a bad deploy is rolled back, no longer kills or re-runs work across the fleet. docs/live-upgrades.md §4.3, minus the forge outbox, reconciliation and late effectful verdicts (separate PRs).

  • Socket activation: the CP takes listen_addr / agent_listen_addr from systemd (LISTEN_FDS, listenfd crate), matched to the configured addresses; falls back to binding. The NixOS module adds krisbuild-control-plane.socket for settings deploys; the service Requires/After it, TimeoutStopSec = 30, stopIfChanged = false (so a switch doesn't also stop the socket and drop queued connections).
  • Graceful stop on SIGTERM/SIGINT: stop accepting, 5 s grace for in-flight requests / the current scheduler pass / merge-queue tick, SSE streams end, agent sockets close with 1012, DB actor drained.
  • Boot lease grant: before the first scheduler pass, lease_until = max(lease_until, now + reconnect_grace_secs) (default 120) for assigned/running instances — CP downtime never counts against a node.
  • Agent reconnect: backoff cap 30 s → 5 s with jitter; immediate redial on 1012; first heartbeat right after Hello; a CP silent for 20 s after having acked is treated as gone.
  • Agent watchdog: connected → unchanged (lease_ttl); disconnected → runs survive disconnected_grace_secs (default 900, ≥ lease_ttl); on reconnect the first ack is authoritative. Safety argument in code and SPEC §7.1.
  • kb follow resubscribes a dropped SSE stream (0/1/2/4/5 s, up to 15 min) instead of reporting a running graph as final; logs de-duplicated by offset.
  • Docs: SPEC §7.1 "Control-plane downtime is not the node's"; forgejo-setup.md "Restarts and deploys" + reconnect_grace_secs.

Deploy notes

  • The first switch introducing the socket unit restarts the CP once more (the old process holds the port). If the socket fails with "address in use": systemctl restart krisbuild-control-plane.socket krisbuild-control-plane.service.
  • No protocol change; any CP/agent version mix keeps today's behaviour. Agents get the new reconnect/grace behaviour once they run this build.

Tests

  • kb-agent: disconnected grace then abort; first ack authoritative on reconnect; 1012 reported as restart; jittered backoff.
  • kb-control-plane: boot lease extension never shortens; tests/restart.rs — graceful stop sends 1012, and after downtime > lease_ttl on the same DB the task completes on attempt 1.
  • kb-cli: resubscription replays no log; backoff; a_control_plane_restart_mid_task_costs_the_task_nothing (real in-process agent + follower, CP down 11 s vs 8 s lease_ttl).
  • flake nixos-module check: socket listenStreams, wantedBy, service requires/after, stopIfChanged, TimeoutStopSec, no socket for configFile.

🤖 Generated with Claude Code

## What Restarting the control plane (a `nixos-rebuild switch`, a crash), or having it down for minutes while a bad deploy is rolled back, no longer kills or re-runs work across the fleet. docs/live-upgrades.md §4.3, minus the forge outbox, reconciliation and late effectful verdicts (separate PRs). - **Socket activation**: the CP takes `listen_addr` / `agent_listen_addr` from systemd (`LISTEN_FDS`, `listenfd` crate), matched to the configured addresses; falls back to binding. The NixOS module adds `krisbuild-control-plane.socket` for `settings` deploys; the service `Requires`/`After` it, `TimeoutStopSec = 30`, `stopIfChanged = false` (so a switch doesn't also stop the socket and drop queued connections). - **Graceful stop** on SIGTERM/SIGINT: stop accepting, 5 s grace for in-flight requests / the current scheduler pass / merge-queue tick, SSE streams end, agent sockets close with **1012**, DB actor drained. - **Boot lease grant**: before the first scheduler pass, `lease_until = max(lease_until, now + reconnect_grace_secs)` (default 120) for assigned/running instances — CP downtime never counts against a node. - **Agent reconnect**: backoff cap 30 s → 5 s with jitter; immediate redial on 1012; first heartbeat right after `Hello`; a CP silent for 20 s after having acked is treated as gone. - **Agent watchdog**: connected → unchanged (`lease_ttl`); disconnected → runs survive `disconnected_grace_secs` (default 900, ≥ `lease_ttl`); on reconnect the first ack is authoritative. Safety argument in code and SPEC §7.1. - **`kb` follow** resubscribes a dropped SSE stream (0/1/2/4/5 s, up to 15 min) instead of reporting a running graph as final; logs de-duplicated by offset. - Docs: SPEC §7.1 "Control-plane downtime is not the node's"; forgejo-setup.md "Restarts and deploys" + `reconnect_grace_secs`. ## Deploy notes - The first switch introducing the socket unit restarts the CP once more (the old process holds the port). If the socket fails with "address in use": `systemctl restart krisbuild-control-plane.socket krisbuild-control-plane.service`. - No protocol change; any CP/agent version mix keeps today's behaviour. Agents get the new reconnect/grace behaviour once they run this build. ## Tests - kb-agent: disconnected grace then abort; first ack authoritative on reconnect; 1012 reported as restart; jittered backoff. - kb-control-plane: boot lease extension never shortens; `tests/restart.rs` — graceful stop sends 1012, and after downtime > `lease_ttl` on the same DB the task completes on attempt 1. - kb-cli: resubscription replays no log; backoff; `a_control_plane_restart_mid_task_costs_the_task_nothing` (real in-process agent + follower, CP down 11 s vs 8 s `lease_ttl`). - flake `nixos-module` check: socket listenStreams, wantedBy, service requires/after, stopIfChanged, TimeoutStopSec, no socket for `configFile`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Control-plane restarts cost running tasks nothing (SPEC §7.1)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by e924dd9245dd0ff9eb87276b57496061dd552427
b09d68e51e
A restart (nixos-rebuild switch, crash) or a control plane down for minutes
(a bad deploy being rolled back) no longer kills or re-runs fleet work:

- Socket activation: the CP takes listen_addr/agent_listen_addr from
  systemd (LISTEN_FDS, via listenfd), matched by address, else binds. The
  NixOS module adds krisbuild-control-plane.socket when `settings` gives
  the addresses; the service requires it, restarts in place
  (stopIfChanged = false keeps the socket open) with TimeoutStopSec = 30.
- Graceful stop on SIGTERM/SIGINT: stop accepting, bounded (5 s) finish of
  in-flight requests, the scheduler pass and merge-queue tick, event
  streams end, agent sockets close with 1012, the db actor drains.
- Boot lease grant: before the first pass, lease_until = max(lease_until,
  now + reconnect_grace_secs) (default 120) for assigned/running instances.
- Agent: redials at once on 1012, else 1-5 s jittered backoff; first
  heartbeat right after Hello; a CP silent for 20 s counts as gone. While
  disconnected, unconfirmed runs survive disconnected_grace_secs (default
  900) instead of lease_ttl; the first LeaseAck of a new connection is
  authoritative.
- kb follow resubscribes to a dropped event stream (backoff, up to 15 min)
  without replaying logs already shown.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CP restarts: semaphore/effectful lease rules, two-stage stop, review fixes
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 63f0e903ca53f98f5b822c09cc8e0ec3f2a03a19
e924dd9245
- Agent lease rules per run: semaphore holders abort at lease_ttl even while
  disconnected (the CP hands their holds on at reap, SPEC §6.4); effectful
  runs without semaphores are never aborted for lease loss (never
  re-queued, SPEC §3.3); everything else keeps the disconnected grace.
- Watchdog reads its allowance before the confirmation stamp.
- Stop in two stages: listeners, scheduler, merge queue and event streams
  first; then 1012 to agents, applying their messages until the close echo.
- A failed listener exits non-zero; LISTEN_FDS is taken before the runtime.
- Agent: silence check has its own arm, websocket writes time out after
  10 s, and 1012 redials at once only once per healthy session.
- kb follow: an outage ends only when a resubscription delivers an event;
  4xx on resubscribe fails at once.
- Docs: socket-unit restarts after an address change, stopping the socket,
  the new lease rules; tests poll instead of tight sleeps.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Strict lease allowance for semaphore holders; hold reruns of lost effectful tasks
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/clippy failed on ares (exit-code)
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test failed on ares (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue the pull-request head changed
63f0e903ca
- A semaphore holder's watchdog allows lease_ttl·3/4 − 1 s, so with its
  lease_ttl/4 tick it fires lease_ttl − 1 s after the last ack arrived:
  before the control plane's lease expires and hands the mutex on. A CP
  outage longer than that still kills such runs (accepted).
- An effectful instance failed as lease-expired may still be running on its
  node; rerunning its graph (API, UI buttons, merge-queue commands) is
  refused with 409 and a named reason until its timeout, counted from its
  start, has passed. New DbError::Conflict; UI buttons get a refusal page.
- SPEC §3.3/§6.4/§7.1 and the operator guide say so; late verdict
  acceptance is noted as the follow-up.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner
@krisbuild r+
Merge origin/main; virtual nodes report no host load; clippy
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/queue queued (#15 in line)
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test failed on ares (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
512a2e1beb
Merge main (verdict integrity, PR #9): SPEC §7.1 keeps both the invalid-
outputs paragraph and the control-plane downtime section.

The example runner's virtual nodes share one host, but each heartbeat
reported the host's /proc memory and load as the node's own `observed` and
`load1`, and admission charges max(reserved, observed) against the node's
declared totals. With the first heartbeat now sent right after Hello, that
reading lands before the first dispatch; on a busy CI host (memory in use
above a virtual node's 16 GB) both `warm` and `cold` looked full and
cache-affinity's e2e never left the queue. Agents now report node-wide
measurements only when the node is the host (`observe_node`; the default
`kb run` fleet), never for a virtual one.

Also: graph::effectful_still_running reads rows into a struct
(clippy::type_complexity).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator

Removed from the merge queue: the pull-request head changed.

Removed from the merge queue: the pull-request head changed.
Author
Owner
@krisbuild r+
Author
Owner
@krisbuild r-
Merge origin/main: single integration-test binary per crate
Some checks failed
krisbuild/queue the pull-request head changed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
02484db281
tests/restart.rs moves into kb-control-plane's tests/it/ as a module; the
restart test in kb-cli's remote.rs and the harness helper followed their
files' renames. flake.nix keeps both main's agent ExecStart binding and the
socket-unit checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner
@krisbuild r+
Collaborator

The pull request head's own build failed (graph 603).

  • nix/dep/cargo-package-listenfd: exit-code on nuxbox, exit 1 — log
  • nix/dep/source: exit-code on nuxbox, exit 1 — log

Push a fix, or comment @krisbuild retry to rerun the failed tasks and requeue.

The pull request head's own build failed ([graph 603](http://192.168.0.2:1337/ui/graphs/603)). - `nix/dep/cargo-package-listenfd`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17420) - `nix/dep/source`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17428) Push a fix, or comment `@krisbuild retry` to rerun the failed tasks and requeue.
Merge origin/main (#20)
Some checks reported errors
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/kb-check superseded
krisbuild/kris/krisbuild/nix/test superseded
krisbuild/kris/krisbuild/nix/clippy superseded
krisbuild/kris/krisbuild/nix/build superseded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: superseded by 733641204219f7619ff9a61bf39091e08d21d390
krisbuild/queue the pull-request head changed
30216c1968
state.rs: AppState keeps main's drvs_present and watch import alongside the
stop tokens and agent task tracker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator

Removed from the merge queue: the pull-request head changed.

Removed from the merge queue: the pull-request head changed.
Author
Owner
@krisbuild r+
Collaborator

The pull request head's own build failed (graph 632).

  • nix/build: exit-code on nuxbox, exit 1 — log
  • nix/clippy: exit-code on nuxbox, exit 1 — log
  • nix/test: exit-code on nuxbox, exit 1 — log

Push a fix, or comment @krisbuild retry to rerun the failed tasks and requeue.

The pull request head's own build failed ([graph 632](http://192.168.0.2:1337/ui/graphs/632)). - `nix/build`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17813) - `nix/clippy`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17818) - `nix/test`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17812) Push a fix, or comment `@krisbuild retry` to rerun the failed tasks and requeue.
Author
Owner

@krisbuild retry

@krisbuild retry
Collaborator

rerunning the failed tasks of graph 632 as graph 649 — the entry waits on it.

rerunning the failed tasks of graph 632 as [graph 649](http://192.168.0.2:1337/ui/graphs/649) — the entry waits on it.
Merge origin/main (#10 enrollment, #8 node health, #22)
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/test dependency-propagated failure
krisbuild/kris/krisbuild/nix/test-deps failed on nuxbox (exit-code)
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check failed on nuxbox (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue head failed: nix/kb-check (exit-code on nuxbox), nix/test-deps (exit-code on nuxbox)
7336412042
- kb-agent conn.rs: the reconnect loop keeps the 1012 immediate redial,
  jitter and 5 s cap, and handles an enrollment refusal first with its own
  jittered backoff capped at 30 s, as before, so a node awaiting approval
  does not poll every 5 s. session() keeps both the enrollment header and
  refusal and the silence/send-timeout handling.
- flake.nix: the nixos-module check keeps the enrolling-node eval and the
  socket-unit assertions.
- kb-cli remote tests: both the restart and the enrollment test; the
  restart test passes spawn_agent's new enroll flag.
- Lease reaping and #8's infra requeue agree: reaping still never requeues
  effectful work, and settle_failure requeues effectful only before the
  payload started, so an Exempt run is never duplicated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator

Removed from the merge queue: the pull-request head changed.

Removed from the merge queue: the pull-request head changed.
Author
Owner
@krisbuild r+
Collaborator

The pull request head's own build failed (graph 650).

  • nix/kb-check: exit-code on nuxbox, exit 1 — log
  • nix/test-deps: exit-code on nuxbox, exit 1 — log

Push a fix, or comment @krisbuild retry to rerun the failed tasks and requeue.

The pull request head's own build failed ([graph 650](http://192.168.0.2:1337/ui/graphs/650)). - `nix/kb-check`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/18101) - `nix/test-deps`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/18106) Push a fix, or comment `@krisbuild retry` to rerun the failed tasks and requeue.
Author
Owner
@krisbuild r+
Merge remote-tracking branch 'origin/main' into lu/cp-restart
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test failed on nuxbox (exit-code)
krisbuild/kris/krisbuild krisbuild kris/krisbuild: one or more tasks failed
krisbuild/queue the pull-request head changed
58ac9b5bfe
Collaborator

The pull request head's own build failed (graph 665).

  • nix/test: exit-code on nuxbox, exit 1 — log

Push a fix, or comment @krisbuild retry to rerun the failed tasks and requeue.

The pull request head's own build failed ([graph 665](http://192.168.0.2:1337/ui/graphs/665)). - `nix/test`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/18276) Push a fix, or comment `@krisbuild retry` to rerun the failed tasks and requeue.
- kb-agent conn.rs: #15's Hello{leases}/Welcome and close-reason refusals
  alongside the 1012 restart, backoff, silence and enrollment handling;
  closed() maps 1012 to Ended::Restarting, a reasoned non-normal close to an
  error, anything else to Ended::Closed.
- Graceful stop covers #14's workers: reconciliation stops with the first
  stage; the outbox makes one final drain in the second, after the last
  scheduler pass and merge-queue tick, and keeps the rest queued.
- graph::rerun takes the config (#14) for the outbox write; the lease-expired
  effectful guard refuses before anything is written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
graph tests: give the lease-expired rerun test its pipeline and graph rows
All checks were successful
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/workspace-deps succeeded
krisbuild/kris/krisbuild/nix/clippy succeeded
krisbuild/kris/krisbuild/nix/test-deps succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/deployed succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue merged
14bdcfcccb
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator

Removed from the merge queue: the pull-request head changed.

Removed from the merge queue: the pull-request head changed.
Author
Owner
@krisbuild r+
krisbuild manually merged commit 3d0651236d into main 2026-09-27 23:27:54 +02:00
Collaborator

Merged as 3d0651236d.

Merged as 3d0651236db1.
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kris/krisbuild!12
No description provided.