Skip to content

fix(dev): OFFICE_ALLOWED_HOSTS now governs the dev server too; render-smoke really escalates a stuck server - #252

Merged
KbWen merged 23 commits into
mainfrom
fix/followups-2026-09-27
Sep 27, 2026
Merged

KbWen merged 23 commits into
mainfrom
fix/followups-2026-09-27

Conversation

@KbWen

@KbWen KbWen commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Two small follow-ups from the 2026-09-26 audit wave, bundled into one PR to save a CI + validator cycle. Each was reviewed on its own branch (fix/dev-allowed-hosts, fix/pack-smoke-teardown).

OFFICE_ALLOWED_HOSTS now also governs the dev server

PR #246 gave the production server a Host allowlist, but npm run dev ignored OFFICE_ALLOWED_HOSTS, so a dev server behind a custom hostname or reverse proxy got Vite's own 403 with no way to allow it. The parser (hostnameFromHeader, parseAllowedHostsEnv) moved to src/utils/hostAllowlist.mjs; server.mjs imports it (behavior unchanged: 480/480 identical responses main vs branch across 6 env configs x 48 Host cases), and vite.config.mjs maps the parsed list to server.allowedHosts only when set. Unset env keeps Vite's default; nothing can become allowedHosts: true.

render-smoke actually escalates a stuck server on POSIX

if (!serverProc.killed) serverProc.kill('SIGKILL') never fired, because .killed turns true when the signal is sent, not when the process exits. It now waits on the real exit event (11 s, just above server.mjs's 10 s drain) before escalating.

Not included: the original pack-smoke "orphan reaper". A fresh review showed the orphaned server.mjs processes were servers agents had started by hand, not a smoke-script bug, and the reaper could kill unrelated processes. It was dropped.

Evidence

  • Combined branch: build clean; npx vitest run 144 files / 2654 tests passed; npm run smoke PASS (4 viewports, 0 errors).
  • Mutation checks: removing the .-entry filter, the entry port-stripping, or the Vite wiring each fails a test.

🤖 Generated with Claude Code

KbWen and others added 23 commits September 26, 2026 17:33
…s, corrupted inference, stale group pose)

A 2026-09-26 audit found three ways an officeLife set-piece could misrepresent real agent
status, the same class ADR-007/008 exist to prevent:

1. Phantom work-claim events: fireWithCast/triggerInteractiveEvent called setActiveEvent
   before a handler's own required-actor check (e.g. deploy-success needs ops) could bail --
   a non-empty but actor-missing cast slipped past the existing empty-cast guard (AVO-191) and
   produced a live banner/confetti/feed entry with nobody performing it. Fixed with a
   REQUIRED_ACTORS map + hasRequiredActors() consulted before setActiveEvent.

2. Idle-gap inference wiped task/label/activeFile/reasonCode/skill to null on every inferred
   tick (buildExtEntry writes u[f] || null, and the inferred update sent {agentId, status}
   only) and minted a fresh changedAt that a work-claim gate could later read as a real recent
   signal. Fixed: idleGapInfer.js carries the agent's own prior externalStatus fields forward;
   store.js keeps changedAt unchanged and skips the new-bubble pop for source:'idle-gap-infer'
   updates.

3. An agent locked in a group event kept performing the group's pose after a real tracked
   status (working/blocked/awaiting-approval/thinking) arrived -- nothing released it, and
   several deferred handler steps (food-delivery, coffee-spill, deploy-success celebrate,
   dog-visit, group-stretch, pm-all-meeting stage 2) re-locked/relocated participants without
   re-checking availability. Fixed: applyExternalStatus releases inGroupEvent/groupTarget on a
   real busy status; the named deferred steps re-verify isAgentAvailable when they run.

Deferred (not implemented, recorded in the spec's Non-goals): tightening deploy-success's
honesty gate to require live status==='done' instead of changedAt-recency -- done is a
10s-transient status by design, so that would make the gate almost never open, and it
conflicts with an existing shipped test contract. Flagged as an open question.

Tests: full suite 135 files / 2533 tests green (vitest run), build clean, bundle-budget PASS
(+1.07%), render-smoke + panel-smoke PASS. New coverage: tests/requiredActorsGate.test.js
(finding #1), tests/groupEventDeferredAvailability.test.js (finding #3 deferred steps),
avo184-equivalence.test.js J2/J2b + K1-K3 (findings #2/#3 at the store.js layer),
idleGapInfer.test.js carry-forward cases (finding #2 at the inference layer).

Spec: docs/specs/honest-office-events.md (new). Amended docs/specs/idle-gap-inference.md and
docs/specs/living-office-events.md with remediation notes -- no cadence/gather-spot/event-list
change (Protected Surfaces untouched).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Re-derived a 2026-09-26 audit's 6 findings against current source and fixed all of them:
- inferStatus.js: a confirmed-unchanged (304) poll response now refreshes the client's
  120s staleness timer, so a long single tool call with no intermediate hook writes no
  longer shows as idle while the session is still genuinely alive (server still confirms
  liveness via its own 300s session window, so a truly dead hook still clears).
- inferStatus.js: SSE now retries every 60s after giving up instead of staying on HTTP
  polling forever; a successful reconnect switches back to the 10s heartbeat cadence.
- store.js: persistedSnapshotKey() excludes the Date.now() _savedAt field from the
  autosave dedup comparison, so localStorage.setItem no longer fires on every 2s tick
  regardless of whether office state changed.
- store.js: salvageStalePersistedState() keeps same-day daily ledgers across a >4h
  tab-closed gap instead of discarding "today done"/"today blocked" along with the
  (intentionally-discarded) stale agent positions.
- desktopNotifier.js: pruneEvictedAgents() now runs every tick, not just on stop, so a
  reused dynamic agent id starts a fresh episode instead of inheriting a stale
  blockedSince timestamp and firing an instant false notification.
- store.js: clearExternalStatus's two eviction sites now prune _storeRecentPicks and
  recurringFailureLog, matching the cleanup already done by applyExternalStatus's
  multi-session path and abortAgentMovement's eviction branch.

Full vitest suite green (2528/2528), build clean, bundle-budget +1.01% (well under +10%
gate), both smoke gates pass. Two fixes verified red-then-green by temporarily reverting
the production line and re-running the new test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… wiring

A fresh adversarial /review of the prior commit (d40af71) returned NOT READY with
6 findings, all now fixed:

- R1 (blocking): connectSSE switched to the 10s heartbeat poller the instant
  startSSEListening returned a cleanup function -- i.e. at retry-attempt-START, not
  connection success -- killing the fast poller on every retry and flapping
  integrationHealth offline on the retry's own transient errors even while GET
  polling was healthy. Fixed via a new onOpen callback wired to SSE's genuine
  network-level 'open' event; only that callback may switch cadence down.
- R2 (blocking): the "no-test rationale" for the SSE retry fix was wrong -- a
  fake-EventSource + stubbed fetch/window + fake-timers harness is feasible in this
  repo's default node test env. Adapted into tests/statusIntegrationSSE.test.js.
- R3: added a wiring test for AC1 (continuous 304 keeps external status alive past
  120s, still under the 300s expiresAt backstop) -- mutation-verified.
- R4: extracted a pure resolvePersisted(raw, now) from loadPersistedState and
  tested it directly -- mutation-verified.
- R5 (doc accuracy): corrected the AC1 mechanism in the spec, the
  shouldRefreshStalenessOnProbe comment, and a stale store.js comment -- the real
  dead-hook bound is the client's own expiresAt (~300s), not server-side
  scanAndMerge change detection.
- R6 (low): added shouldWritePersistedSnapshot() so _savedAt still advances at
  least every 30 minutes even when persisted content is unchanged, so a
  long-quiet-but-open tab doesn't silently lose its 4h freshness grace period.

Full vitest suite green (2539/2539, 134 files), build clean, bundle-budget +1.06%
(under +10% gate), both smoke gates pass. Every fix in this commit was verified
red-then-green by temporarily reverting the specific line/mutation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Cast test gap, busy-reactor clobber

A fresh /review of the honest-office-events remediation (6951ab3) returned NOT READY with 4
blocking items and 2 non-blocking ones. All fixed here.

1. (HIGH, validator gate) docs/specs/honest-office-events.md declared primary_domain but had no
   `## Domain Decisions` section -- added 7 tagged [DECISION]/[TRADEOFF] entries, plus a Target
   Files block inside Acceptance Criteria (the section lint_spec_drift.py actually parses) and
   AC-7/8/9 for the fixes below.

2. (MEDIUM, same honesty class this branch fixes) releasing a participant, or a deferred stage
   abandoning before locking anyone (pm-all-meeting stage 2 when pm was already released), left
   activeEvent live over an empty scene -- reproduced by the reviewer against the real store
   (pm -> working at t=1s, activeEvent still 'pm-all-meeting' at t=13s with nobody in-group; same
   for deploy-success/ops). Fixed: startOfficeLife's existing store.subscribe callback now tracks
   whether the CURRENT activeEvent has ever actually locked a participant and clears it
   (+clearReluctant) the moment the scene goes empty after being non-empty -- guarded against a
   false-positive during dog-visit/group-stretch's async pre-first-lock window.

3. (MEDIUM, AC-1 test coverage) the required-actor check was only exercised via the
   interactive-click path; fireWithCast (the daily/rare-scheduler and real-seed path) has an
   independent call to the same guard, and deleting it there passed the full suite. Added
   tests/fireWithCastRequiredActorsScheduler.test.js (mocked single-event catalog for a
   deterministic daily tick); hand-verified the guard's removal turns it red. Also fixed
   requiredActorsGate.test.js's review-debate negative test, which had been silently isolating the
   pre-existing eventEligible gate instead of the new check, and added the missing group-stretch
   case to groupEventDeferredAvailability.test.js.

4. (MEDIUM, evidence integrity) Gate Evidence timestamps in the Work Log were invented, not
   clock-derived (one was in the future). Corrected via the Drift Log (see Work Log).

5. (LOW) fireInteractionReaction's R1-safe guard only checked agent.inGroupEvent -- a genuinely
   tracked-busy reactor (working/blocked) that simply wasn't locked into a group event could still
   get a fabricated click-reaction bubble stamped over its real voice. Fixed: the guard now also
   checks isAgentAvailable. Decision recorded in the spec: prefer silence over a new machine-side
   BUSY badge (that would be a UI feature addition, out of this state-honesty fix's scope).

6. (owner question, no code change) the reviewer agreed a literal status==='done' gate for
   deploy-success is unworkable; the residual honesty window and a doneAt-based follow-up
   candidate are recorded in the spec's Non-goals for future prioritization.

Tests: full suite 137 files / 2541 tests green. bash .agentcortex/bin/validate.sh -> pass=113
warn=6 fail=0 skip=5 ("Agentic OS integrity check passed"). build/bundle-budget/render-smoke/
panel-smoke all PASS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…isions

Second fresh review returned NOT READY on AC2 (health) plus 2 validator FAILs:

- MEDIUM (blocking): integrationHealth still flapped offline ~0.9s per ~90s SSE
  retry cycle -- by retry time the fast poller can be backed off to its 8s cap,
  so a retry's error burst lands faster than the poller ticks. handleSSEProbe
  now suppresses an SSE-sourced ok:false while the GET-polling channel is active
  and its own last result was ok:true -- polling is authoritative in that
  moment; offline still fires when both channels are genuinely down. Poller
  cadence itself is untouched (no reset-to-base). New health assertion added to
  tests/statusIntegrationSSE.test.js, verified red under the old unconditional
  forwarding.
- MEDIUM (blocking, validator): docs/specs/client-runtime-hygiene.md declared
  primary_domain without a ## Domain Decisions section -- added 4 tagged
  entries. Work Log exceeded the compaction cap (429 lines/27KB) -- compacted
  in place (not moved to archive/), keeping every Gate Evidence receipt in
  original order plus all evidence references.
- LOW: added tests/storePersistenceSeedWiring.test.js covering the
  _lastSavedAtWriteAt load-seed (store.js ~227) via vi.resetModules() + dynamic
  import, since it needs the actual DOM-dependent wiring rather than just the
  pure helper -- verified red under a reverted seed line.

Full vitest suite green (2541/2541, 135 files), build clean, bundle-budget
+1.11% (under +10% gate), both smoke gates pass, validate.sh fail=0
(pass=113 warn=6 skip=5).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… stale-timer races

A fresh /review of the round-2 fix (f8ff3c3) found that the mid-event abandonment auto-clear it
added (finding #4) introduced two new races, both from the same root cause: EVENT_BY_ID's catalog
objects are SHARED (two 'tea-break' fires reuse the same object), so activeEvent identity alone
cannot tell an OLD, already-superseded fire's stale timers apart from a NEW one.

N1: event A's own duration-cleanup timer (captured once, at event.duration) fires unconditionally
and clears a LATER event B that has since taken over, releasing B's cast early. Reproduced: an
ac-broken (A) cleared early at t=1s, tea-break (B) fires at t=2s, A's stale 15000ms timer at
t=15.1s cleared B's activeEvent and released dev/qa 5s early.

N2: dog-visit/group-stretch's staggered per-participant locks keep firing after the event has
been abandoned (scene emptied by an early release), locking the remaining cast under a null
activeEvent -- no mutex, no banner.

Fix: every fired event -- including the ad-hoc lunch-nap time-linked event, which has the same
setActiveEvent-then-later-clear shape -- is assigned a unique, monotonically increasing epoch
(beginEventEpoch). Every deferred handler step and the duration-cleanup timer act ONLY if their
captured epoch is still the live one (isStaleEpoch/endEventEpochIfLive); a stale epoch is a full
no-op that touches nothing, since an agent released early from event A may since belong to a
newer event B. The round-2 abandonment auto-clear now reads/writes this same epoch state instead
of an activeEvent-identity check.

Tests: tests/eventEpochRace.test.js reproduces both N1 and N2 probes exactly and passes; hand
mutation-verified (isStaleEpoch forced to always return false turned both cases red; restoring it
turned both green). Full suite 138 files / 2543 tests green. build/bundle-budget/render-smoke/
panel-smoke all PASS. bash .agentcortex/bin/validate.sh -> pass=111 warn=8 fail=0 skip=5 (the 8
WARNs are pre-existing advisory items, none new).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…test gaps, bubble-clear identity

A fresh /review of the round-3 epoch fix (1031fd6) found one HIGH regression it introduced, plus
two lower-severity gaps -- all fixed here.

F1 (HIGH): fireWithCast never actually verified an event wasn't already active -- every call site
individually checked a (sometimes stale) state snapshot, but the Friday-15:00 time-linked block
calls fireWithCast twice in the SAME tick (tea-break then group-meeting) off one stale snapshot
with no re-check between the two calls. The second call silently superseded the first event's
epoch, and the first event's cast then had no release path left: its own cleanup timer correctly
no-ops as stale per round 3, and it was never part of the second event's cast either -- it stayed
inGroupEvent:true forever with activeEvent:null (frozen; doSchedule/watchdog skip in-group
agents), and permanently disabled the round-2 abandonment auto-clear too. Reproduced: 18/20 runs
stranded on the round-3 HEAD; 0/20 on round-1/round-2 (neither had the epoch mechanism, so a
stale cleanup still unconditionally released its own cast). Fixed: fireWithCast now refuses
(returns false, no side effects) whenever store.getState().activeEvent is already set, read fresh
at call time -- the single-choke mutex the module's own comments always assumed existed. Cadence
is unaffected: the second same-tick event simply never fires.

F2 (test-coverage gap): round-3 tests covered dog-visit's staggered lock and the general cleanup
timer, but not group-stretch's staggered lock (structurally identical) or lunch-nap's cleanup
using its own captured epoch rather than whatever is currently live. Added one test per gap in
tests/eventEpochRace.test.js, each hand mutation-verified.

F3 (LOW): the food-delivery/deploy-success crew reaction-bubble clear timers were also
epoch-gated, so an abandoned event stranded that fabricated bubble indefinitely. Fixed: gate the
clear per-agent on bubble IDENTITY instead (clear only if the agent's bubble is still the exact
value this reaction set).

F4 (Work Log hygiene): added a truthful Guardrails-loaded receipt and an ADR Coverage Check
record (both genuinely performed at round-1 bootstrap, receipts just not written until now, per
the reviewer's explicit request not to backdate) and corrected an Evidence bullet that had
mischaracterized 3 of this Work Log's own validator WARNs as pre-existing.

Tests: tests/eventEpochRace.test.js covers F1 (Friday-15:00 double-fire) and F2 (group-stretch +
lunch-nap epoch coverage); all mutation-verified by hand. Full suite 138 files / 2546 tests green.
build/bundle-budget/render-smoke/panel-smoke all PASS. bash .agentcortex/bin/validate.sh ->
pass=114 warn=5 fail=0 skip=5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y cadence, bubble token

A fresh /review of the round-4 fix (b088b9e) found the F1 regression test didn't actually
discriminate the fix it was meant to prove, and that fix's own restructuring silently starved
Friday's group-meeting cadence. Also addresses two lower-severity follow-ups.

G1 (MEDIUM): the round-4 Friday-15:00 test used Math.random=0.999 for both tea-break's and
group-meeting's random-2-3 cast selection, which happened to pick the SAME cast for both -- so
the test passed identically whether fireWithCast's mutex line was present or deleted. Fixed:
fireWithCast is now exported (matching the existing "exported for tests" precedent already set by
isAgentAvailable/eventEligible/floorTickAllowed) and directly tested with two events that require
DISJOINT casts (eureka->arch only, review-debate->dev+qa) -- an unambiguous, isolated proof of the
mutex line. Hand mutation-verified: removing the mutex line turns the new test red.

G2 (MEDIUM, cadence regression introduced by round 4's own fix): tea-break and group-meeting both
use the random-2-3 cast rule, and the Friday-15:00 time-linked block checks tea-break's condition
first. Once fireWithCast actually had a working mutex, tea-break won every single Friday-15:00
tick, permanently starving group-meeting rather than merely losing one race -- the spec's "cadence
is unaffected" claim from round 4 was false. Fixed per the orchestrator's decision: Friday 15:00
now hands its slot to group-meeting instead of tea-break; tea-break unconditionally keeps 10:00
every day and 15:00 on every other day. New tests assert Friday->group-meeting, Thursday->tea-break.

F3 follow-up (LOW): the round-4 bubble-identity fix compared bubble TEXT alone, but eventBubble's
pools are not guaranteed disjoint -- a phrase can coincidentally repeat across pools or across two
draws from the same pool. Added a per-paint monotonic token, checked alongside the text (not
instead of it, since the text check is what catches a fully external bubble overwrite). New test
forces an identical-text collision via the rng() seam (src/systems/rng.js -- eventBubble routes
through its own seeded seam, not Math.random directly) and proves the newer instance's bubble
survives an older instance's stale clear.

F4 follow-up (LOW): the Drift Log claimed ADR-010 was cited in External References; it was not
(the ADR genuinely does cover store.js -- the fact was right, the citation was simply missing).
Added the row.

Work Log compaction: per .agent/workflows/handoff.md §6, older rounds' Review Feedback / Red Team
Findings / Security Findings (rounds 1-3, already resolved) were moved WHOLE -- byte-identical,
never reworded -- to .agentcortex/context/archive/work/fix-honest-office-events-20260926-part1.md,
with a one-line pointer left in the active log. Protected sections (Gate Evidence, Evidence,
Session Info, etc.) stayed untouched per that same rule.

Tests: full suite 138 files / 2549 tests green (npx vitest run --testTimeout=30000; a handful of
unrelated files hit the default 5s timeout under this session's heavy transform/import load and
pass individually in under 400ms -- isolated as environmental, not a regression). build/
bundle-budget/render-smoke/panel-smoke all PASS. bash .agentcortex/bin/validate.sh -> pass=114
warn=5 fail=0 skip=5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…+ bubble-identity mutant tests

Round-5 review PASS with 4 LOW pre-ship items; this closes all four:

- docs/specs/honest-office-events.md:43 and docs/specs/living-office-events.md:371 no longer
  claim "no change to event cadence" -- both now describe the round-5 (G2) decision that Friday
  15:00 fires group-meeting instead of tea-break.
- docs/specs/honest-office-events.md:321 no longer claims the epoch counter is exported --
  only fireWithCast itself is; the counter/state stays module-private.
- The spec's Files section no longer maps AC-12 to the replaced (non-discriminating) Friday
  test; it now points to the round-5 G1 test that actually covers it.
- tests/eventEpochRace.test.js gains two committed tests: an M5 regression (the crew-reaction
  bubble clear must not depend on epoch liveness -- it still fires after the event is
  abandoned) and a stale-clear-vs-real-bubble test (a stale reaction clear must not wipe a
  real status bubble that has since replaced it). Both hand mutation-verified: re-adding
  isStaleEpoch(epoch) to the clear condition turns the first red; dropping the
  `bubble === reactionBubble` identity check turns the second red. Both restored cleanly (diff
  against pre-mutation source is empty).

Evidence: npx vitest run --testTimeout=30000 -> 138 files / 2551 tests green. npm run build clean
(502.53 kB, bundle-budget +1.21%). smoke + smoke:panel PASS. validate.sh -> pass=114 warn=6
fail=0 skip=5 (warn+1 is the expected transient lock phase=implement vs worklog Current
Phase=review mid-round).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Move r3-r6 review/red-team/security findings and superseded Phase
Summary/Task Description detail verbatim to the existing overflow
archive during /test phase entry compaction (Work Log was at the
worklog.max_kb threshold before the test+handoff receipts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pack-smoke.mjs spawns bin/cli.js, which itself spawns vite as a grandchild
(wrapper pattern). killTree() already reaches that grandchild correctly
from inside the script, but if the harness/OS kills pack-smoke.mjs's own
node process from outside before it reaches cleanup() (a CI/tool timeout,
an operator force-kill), no JS in that process runs again -- finally,
exit, and signal handlers included -- so the cli.js+vite subtree is
orphaned with no self-timeout of its own, for hours. Add a cross-run
stale-PID marker (<tmpdir>/avo-pack-smoke-devchild.pid) written when the
dev-server child spawns and cleared on normal cleanup, so the next
smoke:pack run reaps a leftover subtree instead of leaving it to
accumulate. Also add explicit SIGINT/SIGTERM handlers running full async
cleanup.

render-smoke.mjs had a separate dead-code bug: `if (!serverProc.killed)
serverProc.kill('SIGKILL')` -- `.killed` flips true the instant the
signal is sent, not when the process actually exits, so the SIGKILL
escalation never fired, and the fixed 3s wait was shorter than
server.mjs's own up-to-10s graceful-drain cap. Wait for the real exit
event (11s deadline) before escalating.

panel-features-verify.mjs already tears down correctly (close-event-based
wait); left unchanged after verifying it doesn't leak.

Verified: 4x smoke:pack (2 pre-fix baseline + 2 post-fix, all clean, no
leftover process/port), a seeded stale-marker reap test (dummy process
killed and logged), 2x render-smoke post-fix (clean), 1x smoke:panel
(clean, unchanged), full vitest (139 files / 2579 tests green).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Integrates #246 (server crash hardening), #247 (Codex hook isolation),
#248 (dev-server status parity), #249 (hook robustness/privacy).

Conflict: docs/specs/codex-status-parity-and-done-count.md — both the
Codex filename-namespace-isolation addendum (origin/main) and this
branch's 2026-09-26 hygiene update section were kept in full, Codex
addendum first, hygiene update appended after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…moke fix

Review disproved the premise: hard-killing pack-smoke mid-Assertion-4 and
render-smoke orphaned nothing. The observed orphans (node server.mjs
--port=530x --no-open, parent = a dead bash wrapper) were servers agents
launched manually from their own Bash tool for live probes and never
stopped -- not a smoke-script teardown bug.

The added cross-run PID marker was also unsafe on its own terms: it
taskkill /T /F's whatever live PID is in one shared tmpdir marker with no
identity check, so a concurrent pack-smoke run (or any unrelated process
that happens to land on that PID) can be reaped by another run's cleanup.

Restore scripts/pack-smoke.mjs to origin/main content. Keep
scripts/render-smoke.mjs's exit-event-based teardown fix (the .killed
dead-code bug is real and independently verified correct).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ck too

PR #246 added a Host-header allowlist (OFFICE_ALLOWED_HOSTS) to production
server.mjs, but the Vite dev server never set server.allowedHosts, so the
env var had no effect in dev and a custom-hostname/reverse-proxy dev caller
always got Vite's default 403.

Extract the pure env-parsing step of server.mjs's isAllowedHost
(hostnameFromHeader + a new parseAllowedHostsEnv) into
src/utils/hostAllowlist.mjs, shared by both transports. server.mjs now
imports it instead of duplicating it (behavior-preserving). vite.config.mjs
imports it too and maps OFFICE_ALLOWED_HOSTS into Vite's own
server.allowedHosts, only when non-empty, so an unset env var leaves Vite's
default unchanged and allowedHosts: true is never produced.

Tests: tests/hostAllowlist.test.js (14 unit tests, red before the module
existed) and tests/viteAllowedHosts.test.js (7 behavioral tests via Vite's
real createServer(), raw-TCP Host-header requests since fetch cannot
override Host; 3 of 7 independently re-confirmed red with the wiring
temporarily reverted). Docs updated: README.md, docs/deployment/
DEPLOYMENT.md, docs/specs/engineering-audit-remediation.md (new dated wave).

Evidence: npm run build clean; npx vitest run --testTimeout=30000 - 141
files / 2600 tests green; SMOKE_PORT=5901 npm run smoke PASS (4 viewports,
0 errors); bash .agentcortex/bin/validate.sh - pass=114 warn=6 fail=0
skip=5 (all WARNs pre-existing historical-archive advisories, unrelated to
this branch).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fresh-review follow-up on fix/dev-allowed-hosts (review PASS, 480/480 prod
parity; 1 MEDIUM + 3 LOW findings):

- server.mjs: the comment above ALLOWED_HOSTS_ENV claimed Vite's
  leading-dot wildcard "differs slightly" from AVO's. Verified against
  Vite's own isHostAllowedInternal (node_modules/vite/dist/node/chunks/
  node.js) that the rule is identical; reworded to state the parity
  (that's exactly why vite.config.mjs can reuse the shared parser with no
  translation step).
- tests/viteAllowedHosts.test.js: bootDevServer now wraps createServer()/
  listen() in try/finally so the overridden process.env is restored even
  if either throws, instead of only on the success path.

The Gate Evidence invented-timestamp finding (MEDIUM) and the case-
sensitivity divergence (LOW, accepted/documented) are corrected/recorded
in the (gitignored) Work Log, not this commit.

Verified: node --check server.mjs OK; npx vitest run
tests/viteAllowedHosts.test.js tests/hostAllowlist.test.js --testTimeout=30000
-> 2 files / 21 tests passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ship History + Spec Index updated for fix/client-runtime-hygiene (PR #250):
rotated the oldest Ship History entry (chore-release-v1.6.9) and the oldest
live Spec Index line (subagent-helper-huddle.md) into their archive sections
to hold the 10/30 caps; Domain Decisions consolidated into
docs/architecture/office-runtime.log.md; spec status frozen -> shipped;
Work Log archived to .agentcortex/context/archive/; INDEX.jsonl chain
extended via append_chain_entry.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… fix/honest-office-events

# Conflicts:
#	src/systems/store.js
Consolidates the ship phase for fix/honest-office-events (PR #251): Ship
History + Spec Index entries added and rotated (caps verified), Work Log
archived, INDEX.jsonl chain-appended and verified.

Corrects docs/specs/honest-office-events.md primary_domain (frontend ->
office-runtime, orchestrator-confirmed) and consolidates its Domain
Decisions into docs/architecture/office-runtime.log.md; living-office-events.md
left untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ship closure for PR #252 (fix/dev-allowed-hosts + fix/pack-smoke-teardown,
both quick-win, reviewed separately): Ship History entry at top, oldest
entry (AVO-193) rotated verbatim into ship-history-2026.md, Spec Index note
for engineering-audit-remediation.md extended, both Work Logs archived,
INDEX.jsonl chain extended twice (check_audit_chain: intact), seq 141 -> 142.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@KbWen
KbWen merged commit 879382d into main Sep 27, 2026
8 checks passed
@KbWen
KbWen deleted the fix/followups-2026-09-27 branch September 27, 2026 07:10
KbWen added a commit that referenced this pull request Sep 27, 2026
…ning, and the server stops falling over (#253)

Cuts the 13 PRs merged since v1.6.9 (#240-#252): package.json and both
package-lock.json root version fields 1.6.9 -> 1.6.10, the CHANGELOG
narrative (leading with the OFFICE_ALLOWED_HOSTS upgrade note), and the
release Ship History entry (oldest entry rotated to the archive; SSoT
sequence 142 -> 143). No app code.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant