Node health: report /nix/store free disk in the heartbeat, and floor the lower #35
Loading…
Reference in a new issue
No description provided.
Delete branch "nixstore-free-disk"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What. A node with the
nixfeature now also reports free space on the filesystem that holds /nix/store, asstore_free_disk_mbin its heartbeat. The field is optional in kb-core'sHeartbeat(serde(default, skip_serializing_if = "Option::is_none")). The heartbeat is not hashed, so there is nokbN-bump.When the store is reported:
How the control plane uses it:
min_free_disk_mbfloor applies to the lower of the work dir and the store.heartbeat, and itsunschedulablereason names the filesystem.Docs: SPEC §6.2 (heartbeat listing), §6.3 (pseudocode
min(free_disk, store_free_disk) ≥ floorand the "Taking new work" text) and §7.1 (heartbeat line) are updated, along with themin_free_disk_mbdoc comment, the forgejo-setup config comment and docs/TODO.md (entry struck through as done).Why.
free_disk_mbmeasured only the work dir. If /nix/store was on another filesystem and filled up, nothing showed it until builds failed and the circuit breaker opened. This was the gap left open in the node-health TODO after the 2026-09-27 full-disk incident.Tests.
heartbeat_store_free_disk_is_optional: a heartbeat from an old agent (no field) still parses; when the field is None it is left out of the JSON; a heartbeat with it round-trips.the_store_is_reported_only_when_it_is_a_filesystem_of_its_own(tempdir): store on the same filesystem as the work dir gives None; nonixfeature gives None; a store on another device (procfs) is reported.a_store_that_cannot_be_measured_is_not_reported_full: an unmeasurable store gives None, while the work dir still reads 0.nixreports no store figure.free_disk_floor_judges_the_fuller_filesystem: the floor uses the lower figure and the reason names the filesystem.free_disk_flooris adapted to the new gate signature.node_health::a_node_low_on_store_disk_gets_no_work: a node with plenty of room on its work dir but 100 MB on its store gets no work./api/nodesnames /nix/store, and/ui/nodesshows "100000 MB work dir, 100 MB /nix/store". The node gets work again once the store has room.cargo clippy --workspace --all-targets -- -D warningsandcargo test --workspacepass. Thenix build .#ci.*gate was not run locally (no nix daemon in the dev VM). No files were added, so the filesets are unchanged.Deploy notes. Either side can be upgraded first:
The store check takes effect once both the control plane and an agent whose store is on its own filesystem are upgraded. A node whose /nix/store filesystem is already below
min_free_disk_mb(default 4096) stops taking new work at that point. Checkdf /nix/storeon nuxbox and ares before switching.An agent with the `nix` feature now also reports the free space on the filesystem holding /nix/store (`store_free_disk_mb`, optional in kb-core's Heartbeat). The control plane's `min_free_disk_mb` floor judges the lower of the work dir and the store, and the reason names which one is short ("low disk on /nix/store (…)"). /ui/nodes shows both; GET /api/nodes carries the field in the heartbeat as reported. A full store on its own filesystem was invisible until builds failed and the breaker caught it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>Automated review (reviewer subagent) of
4c7353daeeVerdict: approve, with one SPEC fix before merge.
Gates, run on the commit:
cargo clippy --workspace --all-targets -j2 -- -D warningsclean;cargo test --workspace -j2all pass.nix build .#ci.*not run (no nix in the VM; no files added).Correctness: every free-disk decision goes through
least_free:health::blocked_nodes(scheduler and queue view),load_nodesfor/api/nodes, and the /ui/nodes badge. Mutation checks: reverting eitherblocked_nodesorhttp.rsto the work dir alone fails the new integration test. Compatibility: the field is additive, usesserde(default, skip_serializing_if), and is not hashed, so there is nokbN-bump.Findings:
free_disk ≥ floor. Change it tomin(free_disk, store_free_disk) ≥ floor.free_disk_mbreturns 0 on a failed statvfs, so a failed probe of the store reads as a full store and holds the node. Suggest a falliblefree_disk(path) -> Option<u64>that maps an error toNonefor the store.metadata().dev()of the work dir and/nix/storeon the agent, and sendingNonewhen they match.store_free_disk_mb(features, store: &Path)helper tested with a tempdir. The /ui/nodes "X MB work dir, Y MB /nix/store" format has no assertion.None).df /nix/storeon nuxbox and ares before switching. A node gated for low disk is not logged; a one-timewarn!could be a follow-up.View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.