Priority: low. Nothing is broken without this. Storage writes are best effort by design, and recovery/reconcile heal transient failures with no task loss and no double execution; the cost of a blip is a stale status until the next recovery or reconcile pass picks the row up. A retry policy tightens that window for live processes. It is a refinement, not a gap.
Why there is no transient retry today
EF's EnableRetryOnFailure was never a usable shortcut here, for two reasons:
- Several write paths open explicit transactions, which EF's retrying execution strategy rejects outright.
- More fundamentally, its classifier treats a command timeout as transient and replays the whole operation, and a timeout can arrive after the server already committed. Blind replay is unsafe in this codebase:
- The CAS advance does not consume its own token.
usp_UpdateCurrentRunCas guards on ScheduleVersion = @ExpectedScheduleVersion, then increments CurrentRunCount without changing that version (the Postgres CTE has the same shape). After an ambiguous commit the identical retry passes the guard again: one run counted twice, MaxRuns spent early, a series can end before its time.
SetStatus at audit Full is not idempotent: each execution inserts a fresh StatusAudit row and re-stamps LastExecutionUtc.
- A PK violation on a
Persist retry proves prior success only when it is exactly PK_QueuedTasks. A TaskKey collision or the (ParentTaskId, ScheduledExecutionUtc) occurrence index firing is a different event and must not be read as "already inserted".
The policy to build (per operation, never a global flag)
- Failure before the command was sent: bounded retry, always safe.
Persist ambiguous: read back by PK, verify the row's identity, then decide.
- Status write ambiguous: read the row; retry only if the desired state is not already there.
- CAS advance: automatic retry only if the operation is changed to consume its token (bump
ScheduleVersion in the same statement); until then, verify-then-retry.
- A caller-requested cancellation never turns into a retry.
- Outage behavior: few attempts, exponential backoff with jitter, and a global budget or circuit breaker so a recovering database is not flooded by its own clients.
Gate before enabling anything
Fault injection before, during and after commit on every write shape, plus a sustained-outage run proving the budget holds and recovery still converges. EF's own guidance treats a commit failure as unknown state and recommends verification over replay; that is the bar.
Priority: low. Nothing is broken without this. Storage writes are best effort by design, and recovery/reconcile heal transient failures with no task loss and no double execution; the cost of a blip is a stale status until the next recovery or reconcile pass picks the row up. A retry policy tightens that window for live processes. It is a refinement, not a gap.
Why there is no transient retry today
EF's
EnableRetryOnFailurewas never a usable shortcut here, for two reasons:usp_UpdateCurrentRunCasguards onScheduleVersion = @ExpectedScheduleVersion, then incrementsCurrentRunCountwithout changing that version (the Postgres CTE has the same shape). After an ambiguous commit the identical retry passes the guard again: one run counted twice,MaxRunsspent early, a series can end before its time.SetStatusat audit Full is not idempotent: each execution inserts a freshStatusAuditrow and re-stampsLastExecutionUtc.Persistretry proves prior success only when it is exactlyPK_QueuedTasks. ATaskKeycollision or the(ParentTaskId, ScheduledExecutionUtc)occurrence index firing is a different event and must not be read as "already inserted".The policy to build (per operation, never a global flag)
Persistambiguous: read back by PK, verify the row's identity, then decide.ScheduleVersionin the same statement); until then, verify-then-retry.Gate before enabling anything
Fault injection before, during and after commit on every write shape, plus a sustained-outage run proving the budget holds and recovery still converges. EF's own guidance treats a commit failure as unknown state and recommends verification over replay; that is the bar.