forge: durable outbox and webhook reconciliation (SPEC §8.3) #14
Loading…
Reference in a new issue
No description provided.
Delete branch "lu/forge-outbox"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What
docs/live-upgrades.md §4.3, "Forge outbox" and "Webhook reconciliation".
Forge outbox. Every forge write — commit statuses, merge-queue comments, mark-merged — is a row in
forge_outbox, written in the same transaction as the state change that causes it (graph created →pending; graph terminal → aggregate + final per-task statuses; each merge-queue transition →krisbuild/queuestatus, comment, mark-merged). One worker drains it with exponential backoff persisted in the row (2 s doubling to 5 min); dropped on a 4xx other than 408/429 or after a day of failures; everything due is retried on restart. Statuses coalesce per (repo, sha, context) — the newest unsent one replaces the older. Comments are deleted as soon as the forge accepts them (a crash in between may post one twice — accepted, documented); writes to one PR stay ordered. The in-memory reporter and the driver's direct comment/merge calls are gone (ForgeWriter, single attempt,forgejo/writer.rs).Reconciliation. Push and PR webhooks record their head per (repo, ref) in
forge_refs. On boot and everyforge_reconcile_secs(default 300,0disables), each repo's branches and open PRs are listed; a differing head goes through the webhook's own handlers (http::on_push/on_pull_request, refactored to be shared) withtrigger.source = "reconcile"— all filters apply. A recorded PR no longer open is fed through asclosed. A repo's first run only seeds. For merge-queue repos, new comments are fed to the same command handler, which claims each comment id inforge_commentsbefore applying it, so a command seen by webhook and by listing is applied once.Schema: additive only (
forge_outbox,forge_refs,forge_reconcile,forge_comments); older binaries still run on it.Why
A CP stop at the wrong moment left a check
pendingforever, and webhooks delivered during a restart were lost (Forgejo doesn't redeliver).Docs
SPEC §8.3 (new), §8 adapter surface, §9 schema; forgejo-setup.md:
forge_reconcile_secs, "Restarts and missed deliveries".Tests
tests/forge_outbox.rs: graph created and finished while the forge returns 503, server stopped; a new server on the same DB posts onlysuccess.tests/reconcile.rs: missed push built once (reconcile); already-seen heads andkb-queue/*skipped; missed PR update built, missed close cancels; missedr+applied, a late webhook for it is "already handled", edited old comments ignored.tests/forge_api.rs: writer calls and listings. Fake forge gained listings and a write-failure switch.🤖 Generated with Claude Code
@krisbuild r+
Removed from the merge queue: the pull-request head changed.
@krisbuild r+
@krisbuild r-
@krisbuild r+
The pull request head's own build failed (graph 611).
nix/kb-check: exit-code on nuxbox, exit 1 — logPush a fix, or comment
@krisbuild retryto rerun the failed tasks and requeue.Head CI failed only because
nix/kb-check's .drv was missing from attic (the eval-push race, fixed in the deployment since). Pushed the missing drvs and reran the failed tasks: graph 646 is green.@krisbuild retry
Removed from the merge queue: the merge conflicts in SPEC.md, crates/kb-control-plane/src/db.rs, crates/kb-control-plane/src/scheduler/status.rs, crates/kb-control-plane/src/state.rs.
Conflicts: SPEC.md and db.rs (agent_enrollments beside the forge tables; both kept), state.rs and scheduler/status.rs. The node-health status behaviour — infra vs verdict descriptions, a requeued attempt re-posting `pending` ("... (retrying on another node)") — now runs through the outbox sweep: the per-task bookkeeping is keyed on (state, attempt), and completion enqueues the final statuses in its own transaction as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>@krisbuild r+
Merged as
3e02404503.