Skip to content

perf(agent-manager): Codex discovery reads whole session files to parse only the first line #258

Description

@codeaholicguy

Problem

CodexSessionLocator.discoverSessionFilesInDateDirs (packages/agent-manager/src/providers/codex/CodexSessionLocator.ts) runs for Codex processes that can't be matched by resume ID, the session mapping, or the registry cache.

  • It reads every session file in a ±1-day window around each process start in full, only to parse the first line (session_meta).
  • It keeps every file's full contents in contentCache at the same time.
  • It repeats this on every refresh for processes that never match, so the cost never goes away. Non-agent codex processes are one such case (see fix(agent-manager): non-agent codex helper processes (app-server, sandbox) are listed as running agents #259).
  • Because windows are derived per process start date, long-running processes spread the scan across many days.

Evidence

  • On a machine with 21 Codex processes, discovery read 38 MB per refresh (17 files) just to get first lines.
  • contentCache holds all of those strings at once, which raises peak RSS.

Proposed approach

  • Read only a bounded head of each file (e.g. fs.readSync of up to 64 KiB, stopping at the first newline) to get session_meta.
  • Remove contentCache, or limit it to head bytes. Legacy matches then read through the incremental summary utility (perf(agent-manager): shared incremental session-summary cache (stop full transcript re-parse on every refresh) #256).
  • Cache the metadata extracted from each file's head, keyed by path, inode and size: session_meta doesn't change once written.
  • Cache negative results: (pid, startTime) → "no session", re-checked only when the date directories change (directory mtime), with a periodic re-check of at most every 30 s.

Acceptance criteria

  • Discovery reads ≤ 64 KiB per candidate file; a test shows a 100 MB fixture file costs ≤ 64 KiB.
  • No full file contents are kept in contentCache or elsewhere by discovery.
  • After the first refresh, a refresh with unchanged date directories reads 0 bytes from discovery candidates.
  • An unmatched process whose date directories are unchanged is not rescanned (negative cache test); a new file in its window triggers a rescan.
  • Existing legacy/birthtime matching tests pass with identical results.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions