Repository navigation
[agent-job-health] Agent Job Health: 2026-10-02 fleet failure rate 11.8% (25/212), novel Avenger + redact-secrets clusters #65133
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-10-05T23:53:59.883Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Summary
--start-date -1d --count 3000), but ran into its own count/time ceiling before reaching back to 2026-10-01 23:34 UTC — so this report covers the most recent ~15h of the 24h window, not the full 24h. Treat the rates below as representative of the busy tail of the day, not a strict 24h figure.in_progressat analysis time — this very Agent Job Health Monitor run — excluded from rate math)skippedagent jobs are excluded per scope)Warning
The median-vs-mean gap (0.0% vs 16.3%) confirms the fleet rate is driven by a handful of chronic/novel offenders, not a broad regression. One workflow (Avenger) failed its agent job on 6/6 observed runs (100%), and a distinct "Redact secrets in logs" failure recurs across 6 unrelated workflows — novel and not matched to any open issue found in this session's search (filed as a new issue, see below).
Failure Rate by Step
Tracked Failures
Caveat on "tracked": every prior recorded window (2026-08-23, 2026-08-28, 2026-09-23) shows this exact "Execute GitHub Copilot CLI step failure" signature as a recurring top cluster, previously linked to issues #54186 / #55413 in earlier windows.
search_issuesin this session returned 0 directly visible matches for "Copilot CLI segfault"-style queries, with several candidate results filtered out by integrity policy before I could read them to confirm open/closed status. I'm classifying this cluster as tracked by historical precedent, not by a freshly confirmed open issue — please verify #54186/#55413 (or their successors) are still open and still match before assuming this is fully covered.Novel Failure Clusters
1. Avenger — "Execute Codex CLI" step failure (6 runs, 1 workflow, 100% of Avenger's observed runs)
search_issuesfound nothing visibly open for "avenger codex" (candidates were integrity-filtered).2. "Redact secrets in logs" step failure across 6 unrelated workflows (6 runs, 6 distinct workflows) — issue filed
<engine>CLI") is recorded assuccess— only the shared "Redact secrets in logs" post-step fails. This points at a shared pipeline component (the log-redaction step itself), not at any individual workflow's prompt/logic.Schedule Heartbeat
The fleet has ~230 active schedule-triggered workflows (from
gh-aw status). A full per-workflowlist_workflow_runssweep across all 230 was out of scope for this run's budget. Instead, sub-daily (hourly-or-more-frequent) workflows were cross-checked against the observed-run sample; the one apparent gap (Daily Trajectory Grader Implementer, "every 30 minutes") was individually verified via the GitHub Actions API and found to have run at 2026-10-02 23:31:02 UTC — just outside the edge of this session's observed sample window, not an actual blind spot.No blind spots detected among the sub-daily-cadence workflows checked. A full sweep of all ~230 schedule-triggered workflows (including daily/weekly ones) was not completed this run — recommend a follow-up pass with a wider log-fetch budget if deeper coverage is wanted.
View per-workflow breakdown (workflows with ≥1 agent-job failure)
Recommendations
References:
All reactions