Node health: requeue infra failures, quarantine failing nodes, infra vs verdict statuses #8
Loading…
Reference in a new issue
No description provided.
Delete branch "node-health"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Node health, from the 2026-09-27 ares outage (disk full → every job failed with
spawn-errorfor ~9h while ares kept getting work).scheduler/health.rs::classify, new SPEC §3.3.1). Unknown classes are verdicts (never retried, never held against a node). kb-agent:spawn-errornow strictly means the payload never started; a failure waiting on a started process is the newwait-error.lostwith its exit kept, a new attempt avoids that node (hard exclusion only while another connected node could take it). Pre-payload failures requeue for any effect class, post-payload only for pure/idempotent. Budgetinfra_retries(default 2). Failed-run usage is dropped so it doesn't skew learned requests.min_free_disk_mb(default 4096), and a circuit breaker afterquarantine_after(3) consecutive node-attributed infra failures with aquarantine_secs(600) cool-off, then probation (one task) → one success closes it. Derived from instance history each pass; never touchesnodes.state, so manual drains are unaffected. Shown on/ui/nodes,/api/nodesand the queue's "why not scheduled".failed: exit 1 on aresvsinfra: spawn-error on ares, and… (retrying on another node)while a requeue is pending. Per-task statuses dedupe by (state, attempt).No new tables/columns; one additive index. docs/forgejo-setup.md documents the four config keys; docs/TODO.md has the open items.
Review notes / follow-ups:
git-checkoutcounts against the node, so a rev uncheckoutable everywhere can quarantine all nodes for one cool-off (documented, bounded).free_disk_mbmeasures the work dir only, not/nix/store.🤖 Generated with Claude Code
@krisbuild r+
Merge candidate failed: candidate graph failed
@krisbuild retry
The candidate failure was infra:
nix/buildwas placed on nuxbox, but its.drv(evaluated on ares) was never pushed to attic. The closure has been pushed by hand; see the PR thread for the cache-push follow-up.Removed from the merge queue: the merge conflicts in crates/kb-control-plane/src/ws.rs, crates/kb-core/src/protocol.rs.
Merged main in (conflicts in ws.rs / protocol.rs resolved; main's new
expansion-invalid/outputs-invalidclasses classified as verdicts) and added thedrv-fetchinfra class for an unfetchable.drv. Reviewed.@krisbuild r+
@krisbuild r-
Yielding the queue slot to #18 (faster test suite); will merge main into this branch and re-approve once it lands.
Merged main (#18) in;
node_healthintegration tests moved into the singletests/itbinary. Reviewed.@krisbuild r+
The merge candidate's build failed (graph 625).
nix/kb-check: exit-code on nuxbox, exit 1 — logComment
@krisbuild retryto requeue with a fresh candidate on the current base.@krisbuild retry
Infra again:
nix/kb-checklanded on nuxbox, its.drv(evaluated on ares) was not in attic. Pushed the closure by hand; the lasting fix is agent cache push (pixienixkrisbuild-agent-push).Merged as
fdbb9c5aa9.