Conversation
Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
The sweeper that expires leases and fires delegation deadlines was spawned fire-and-forget: its JoinHandle was only aborted at shutdown, so a panic killed the task while the CP kept serving — leases never expired and deadlines never fired, with /health still answering ok (openabdev#1474). - run_sweeper arms an exit guard that flips the health signal to dead on ANY task teardown (return, panic unwind, abort) and stamps a heartbeat after each completed pass, so a wedged mid-sweep stall is as visible as a dead task. - /health now reports 200 ok only while the sweeper is beating within a 10s slack; otherwise 503 with the reason (not started / dead / stalled). - main polls the sweeper's JoinHandle against the server future: any termination is fatal — the process exits non-zero so the orchestrator restarts a clean CP rather than recovering a poisoned loop. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Contributor
|
Caution This PR is missing a Discord Discussion URL in the body. All PRs must reference a prior Discord discussion to ensure community alignment before implementation. Please edit the PR description to include a link like: |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #1474
Summary
crates/openab-cp's lease/deadline sweeper was spawned fire-and-forget: itsJoinHandlewas only.abort()ed at shutdown, so a panic insiderun_sweeperkilled the task while the CP kept accepting registrations and delegations — leases never expired, deadlines never fired, and/healthstill answeredok. Same failure class as the session-pool reaper bug fixed in #1457.run_sweepernow arms aSweeperExitGuardfor the task's whole lifetime: itsDropflips the health signal toDeadon every exit — return, panic unwind, abort — so/healthturns over during the task's own teardown, before itsJoinHandleis even observed.Instant), so a sweeper wedged inside a pass — a case a bareJoinHandlewatch can never observe — goes stale on/healthafterSWEEPER_STALL_SLACK(10s, i.e. 10 ticks of silence from a 1s loop).GET /healthreturns200 okonly while the sweeper is beating; otherwise503with the reason in the body (sweeper not started/sweeper dead/sweeper stalled). A CP servingapp(state)without a sweeper reports down from the start rather than lying healthy.mainnow polls the sweeper'sJoinHandleintokio::select!against the server future: any termination is fatal — the process exits non-zero so the orchestrator restarts a clean CP. (The issue accepts fatal or restart-with-backoff; fatal is chosen because a restarted loop would re-trip the same fault.)The branch also carries a cherry-pick of
53ab5ff8(test: guard cargo fmt cleanliness, resolve clippy --all-targets drift, also ondevin/issue-1544/ PR #1552): the base tree is notcargo fmt --check-clean under the current toolchain, so a verified-clean tree requires that workspace-wide reformat + hygiene test. Identical mechanical output merges cleanly whichever PR lands first; the issue fix itself ise034393fonly.Review Contract
Goal
Make sweeper death impossible to miss:
/healthreflects sweeper liveness (never-started, dead, or stalled) and a terminated sweeper fails the process so a supervisor restarts it.Non-goals
/healthchecks (registry depth, router health) — out of scope.cargo fmt --checkworkspace-wide); it is not part of the fix semantics and is identical to the commit already proposed in PR docs(adr): OpenAB Mac Agent — cloud brain (k8s) + thin macOS executor over Tailscale #1552.Accepted Residual Risks
503is possible if/healthis hit between server start and the sweeper's first pass (~ms window;interval's first tick is immediate). Mitigation: readiness probes retry; the signal is correct throughout.SWEEPER_STALL_SLACKis a compile-time constant (10s), not a config knob — deliberate, matches the hardcoded 1s tick.panic = "abort"build the exit guard doesn't run, but the process aborts anyway — still fatal, still observable.Acceptance Criteria
cargo test -p openab-cp— 143 unit + 12 integration tests pass, incl. newsweeper_health_lifecycle,sweeper_exit_guard_marks_dead_during_unwind, andtests/sweeper_health.rs(never-started → 503, alive → 200, dead → 503).cargo test --workspace— full suite green.cargo fmt --all -- --check— clean.cargo clippy --workspace --all-targets -- -D warnings— clean.cargo test -p openab-cp --test sweeper_healthfails on the base (exit 101;/healthansweredokfor never-started and dead sweeper).Follow-ups
SweeperHealthsignal + supervisor hook point already exist;mainis where the policy lives.Test plan
cargo fmt --all -- --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningsGenerated with Devin