Title: [Bug]: Actor's worker IP is not revalidated after a node reboot — stays RUNNING, unreachable
Relatives: #1526 (worker pod disappears → CRASHED, unrecoverable), #1665/#1518/#50 (stuck in a transitional state after a failed lifecycle op). This is the step before either: the actor never leaves RUNNING, silently pointing at a dead address.
What happened
After a full host reboot, worker pods come back Running with the same name but a new pod IP (fresh network sandbox). The actor controller never re-resolves it: kubectl ate get actor / ax describe task keep reporting RUNNING/Ready at the old IP. First sign of trouble is a request failing:
upstream reset: ... delayed connect error: No route to host
(also seen via ax ssh: rpc error: code = Unavailable ... remote connection failure)
Trying to fix it by hand makes it worse — suspend actor fails (runsc checkpoint: exit status 128) and strands the actor in SUSPENDING with no way out (#1665/#50). Only recovery: delete and recreate.
Repro
- Single-node k3s,
--kind profile, WorkerPool with 2 replicas.
- Create an actor (or an AX Task — same mechanism), confirm it's reachable, note its worker IP.
- Reboot the host.
- Worker pod: same name, new IP (
kubectl get pod <worker> -o wide before/after). Actor: still RUNNING at the old IP → unreachable, and suspend strands it in SUSPENDING.
Reproduced identically for a plain counter demo actor and for an AX Task.
Expected
The controller should revalidate/re-resolve a RUNNING actor's worker IP (or treat a recreated sandbox like a disappeared worker, #1526) — not report healthy indefinitely with no health signal.
Environment
Substrate 672533541dbf (main), --kind profile, single-node k3s 1.36.4, also reproduced through AX v0.3.0.
Title: [Bug]: Actor's worker IP is not revalidated after a node reboot — stays RUNNING, unreachable
What happened
After a full host reboot, worker pods come back
Runningwith the same name but a new pod IP (fresh network sandbox). The actor controller never re-resolves it:kubectl ate get actor/ax describe taskkeep reportingRUNNING/Readyat the old IP. First sign of trouble is a request failing:(also seen via
ax ssh:rpc error: code = Unavailable ... remote connection failure)Trying to fix it by hand makes it worse —
suspend actorfails (runsc checkpoint: exit status 128) and strands the actor inSUSPENDINGwith no way out (#1665/#50). Only recovery: delete and recreate.Repro
--kindprofile, WorkerPool with 2 replicas.kubectl get pod <worker> -o widebefore/after). Actor: stillRUNNINGat the old IP → unreachable, andsuspendstrands it inSUSPENDING.Reproduced identically for a plain
counterdemo actor and for an AX Task.Expected
The controller should revalidate/re-resolve a
RUNNINGactor's worker IP (or treat a recreated sandbox like a disappeared worker, #1526) — not report healthy indefinitely with no health signal.Environment
Substrate
672533541dbf(main),--kindprofile, single-node k3s 1.36.4, also reproduced through AX v0.3.0.