Node health: requeue infra failures, quarantine failing nodes, infra vs verdict statuses #8

Manually merged
krisbuild merged 6 commits from node-health into main 2026-09-27 21:29:39 +02:00
Owner

Node health, from the 2026-09-27 ares outage (disk full → every job failed with spawn-error for ~9h while ares kept getting work).

  • Infra vs verdict classification (scheduler/health.rs::classify, new SPEC §3.3.1). Unknown classes are verdicts (never retried, never held against a node). kb-agent: spawn-error now strictly means the payload never started; a failure waiting on a started process is the new wait-error.
  • Requeue infra failures on another node (#8): in the same transaction as the failure event; the failed attempt ends lost with its exit kept, a new attempt avoids that node (hard exclusion only while another connected node could take it). Pre-payload failures requeue for any effect class, post-payload only for pure/idempotent. Budget infra_retries (default 2). Failed-run usage is dropped so it doesn't skew learned requests.
  • Unhealthy nodes take no new work (#7): free-disk floor min_free_disk_mb (default 4096), and a circuit breaker after quarantine_after (3) consecutive node-attributed infra failures with a quarantine_secs (600) cool-off, then probation (one task) → one success closes it. Derived from instance history each pass; never touches nodes.state, so manual drains are unaffected. Shown on /ui/nodes, /api/nodes and the queue's "why not scheduled".
  • Status text (#11): failed: exit 1 on ares vs infra: spawn-error on ares, and … (retrying on another node) while a requeue is pending. Per-task statuses dedupe by (state, attempt).

No new tables/columns; one additive index. docs/forgejo-setup.md documents the four config keys; docs/TODO.md has the open items.

Review notes / follow-ups: git-checkout counts against the node, so a rev uncheckoutable everywhere can quarantine all nodes for one cool-off (documented, bounded). free_disk_mb measures the work dir only, not /nix/store.

🤖 Generated with Claude Code

Node health, from the 2026-09-27 ares outage (disk full → every job failed with `spawn-error` for ~9h while ares kept getting work). - **Infra vs verdict classification** (`scheduler/health.rs::classify`, new SPEC §3.3.1). Unknown classes are verdicts (never retried, never held against a node). kb-agent: `spawn-error` now strictly means the payload never started; a failure waiting on a started process is the new `wait-error`. - **Requeue infra failures on another node (#8)**: in the same transaction as the failure event; the failed attempt ends `lost` with its exit kept, a new attempt avoids that node (hard exclusion only while another connected node could take it). Pre-payload failures requeue for any effect class, post-payload only for pure/idempotent. Budget `infra_retries` (default 2). Failed-run usage is dropped so it doesn't skew learned requests. - **Unhealthy nodes take no new work (#7)**: free-disk floor `min_free_disk_mb` (default 4096), and a circuit breaker after `quarantine_after` (3) consecutive node-attributed infra failures with a `quarantine_secs` (600) cool-off, then probation (one task) → one success closes it. Derived from instance history each pass; never touches `nodes.state`, so manual drains are unaffected. Shown on `/ui/nodes`, `/api/nodes` and the queue's "why not scheduled". - **Status text (#11)**: `failed: exit 1 on ares` vs `infra: spawn-error on ares`, and `… (retrying on another node)` while a requeue is pending. Per-task statuses dedupe by (state, attempt). No new tables/columns; one additive index. docs/forgejo-setup.md documents the four config keys; docs/TODO.md has the open items. Review notes / follow-ups: `git-checkout` counts against the node, so a rev uncheckoutable everywhere can quarantine all nodes for one cool-off (documented, bounded). `free_disk_mb` measures the work dir only, not `/nix/store`. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
proc::run now classifies its own failures: a failed spawn(2) is spawn-error,
a failure waiting on a started child is wait-error (after tearing the tree
down). The control plane relies on spawn-error meaning nothing of the payload
ran to requeue even effectful tasks after an infrastructure failure
(SPEC §3.3.1).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Node health: requeue infra failures elsewhere, disk floor, circuit breaker
Some checks failed
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/build cached
krisbuild/kris/krisbuild/nix/clippy cached
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue the merge conflicts in crates/kb-control-plane/src/ws.rs, crates/kb-core/src/protocol.rs
46693bafd4
One classification (scheduler/health.rs, SPEC §3.3.1) decides whether a
failure class is the task's verdict or an infrastructure failure, and whether
the payload had started:

- An agent-reported infra failure becomes a new attempt in the same
  transaction (the old one ends lost, its exit kept), avoiding every node that
  failed it while another node could take the task, within infra_retries
  (default 2). Pre-payload classes requeue even effectful tasks. The failed
  run's zero usage is not learned from.
- A node reporting less than min_free_disk_mb (default 4096) free, or whose
  last quarantine_after (default 3) outcomes were node-attributed infra
  failures, takes no new work; the quarantine lifts after quarantine_secs
  (default 600) into probation (one task at a time) until a success. Derived
  from instances each pass (new index instances_by_node), never written to
  nodes.state, so a manual drain is untouched. Shown on /ui/nodes,
  GET /api/nodes and the queue's why-not.
- Per-task statuses read 'failed: exit 1 on ares' vs 'infra: spawn-error on
  ares', and a requeued task stays pending as '... (retrying on another
  node)'; a new attempt re-posts its pending status.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner
@krisbuild r+
Collaborator

Merge candidate failed: candidate graph failed

Merge candidate failed: candidate graph failed
Author
Owner

@krisbuild retry

The candidate failure was infra: nix/build was placed on nuxbox, but its .drv (evaluated on ares) was never pushed to attic. The closure has been pushed by hand; see the PR thread for the cache-push follow-up.

@krisbuild retry The candidate failure was infra: `nix/build` was placed on nuxbox, but its `.drv` (evaluated on ares) was never pushed to attic. The closure has been pushed by hand; see the PR thread for the cache-push follow-up.
Collaborator

Removed from the merge queue: the merge conflicts in crates/kb-control-plane/src/ws.rs, crates/kb-core/src/protocol.rs.

Removed from the merge queue: the merge conflicts in crates/kb-control-plane/src/ws.rs, crates/kb-core/src/protocol.rs.
ws.rs: main's rejected-Outputs path (set_exit + failed) alongside the infra
settle_failure path. protocol.rs: both sides' ExitInfo class docs. main's
expansion-invalid / outputs-invalid are verdicts in the SPEC §3.3.1 table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
kb-agent: an unfetchable .drv fails drv-fetch before realise
Some checks are pending
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/clippy cached
krisbuild/kris/krisbuild/nix/build cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue testing: http://192.168.0.2:1337/ui/graphs/594
3e63954dda
A NixDrv whose .drv is neither on the node nor in any substituter used to run
nix build anyway and fail as an exit-code verdict. It is now the infra class
drv-fetch, raised before the payload runs and not held against the node, so
the control plane requeues it where the .drv exists (SPEC §3.3.1; classified
in scheduler/health.rs alongside the merge of main).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner

Merged main in (conflicts in ws.rs / protocol.rs resolved; main's new expansion-invalid/outputs-invalid classes classified as verdicts) and added the drv-fetch infra class for an unfetchable .drv. Reviewed.

@krisbuild r+

Merged main in (conflicts in ws.rs / protocol.rs resolved; main's new `expansion-invalid`/`outputs-invalid` classes classified as verdicts) and added the `drv-fetch` infra class for an unfetchable `.drv`. Reviewed. @krisbuild r+
Author
Owner

@krisbuild r-

Yielding the queue slot to #18 (faster test suite); will merge main into this branch and re-approve once it lands.

@krisbuild r- Yielding the queue slot to #18 (faster test suite); will merge main into this branch and re-approve once it lands.
tests: node_health joins the single tests/it binary
All checks were successful
krisbuild/kris/krisbuild/nix/workspace-deps cached
krisbuild/kris/krisbuild/nix/hello cached
krisbuild/kris/krisbuild/nix/test-deps cached
krisbuild/kris/krisbuild/nix/world cached
krisbuild/kris/krisbuild/nix/clippy cached
krisbuild/kris/krisbuild/nix/test succeeded
krisbuild/kris/krisbuild/nix/build succeeded
krisbuild/kris/krisbuild/nix/kb-check succeeded
krisbuild/kris/krisbuild krisbuild kris/krisbuild: all tasks succeeded
krisbuild/queue merged
d4fb56e386
tests/node_health.rs becomes the tests/it::node_health module (SPEC-free
layout change from #18); its fail_class/heartbeat_disk helpers already sit in
tests/it/common and are used, so nothing needs a dead_code allow.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Owner

Merged main (#18) in; node_health integration tests moved into the single tests/it binary. Reviewed.

@krisbuild r+

Merged main (#18) in; `node_health` integration tests moved into the single `tests/it` binary. Reviewed. @krisbuild r+
Collaborator

The merge candidate's build failed (graph 625).

  • nix/kb-check: exit-code on nuxbox, exit 1 — log

Comment @krisbuild retry to requeue with a fresh candidate on the current base.

The merge candidate's build failed ([graph 625](http://192.168.0.2:1337/ui/graphs/625)). - `nix/kb-check`: exit-code on nuxbox, exit 1 — [log](http://192.168.0.2:1337/ui/instances/17740) Comment `@krisbuild retry` to requeue with a fresh candidate on the current base.
Author
Owner

@krisbuild retry

Infra again: nix/kb-check landed on nuxbox, its .drv (evaluated on ares) was not in attic. Pushed the closure by hand; the lasting fix is agent cache push (pixienix krisbuild-agent-push).

@krisbuild retry Infra again: `nix/kb-check` landed on nuxbox, its `.drv` (evaluated on ares) was not in attic. Pushed the closure by hand; the lasting fix is agent cache push (pixienix `krisbuild-agent-push`).
krisbuild manually merged commit fdbb9c5aa9 into main 2026-09-27 21:29:39 +02:00
Collaborator

Merged as fdbb9c5aa9.

Merged as fdbb9c5aa986.
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kris/krisbuild!8
No description provided.