Skip to content

storage: opt-in transient-retry policy, classified per operation #56

Description

@GiampaoloGabba

Priority: low. Nothing is broken without this. Storage writes are best effort by design, and recovery/reconcile heal transient failures with no task loss and no double execution; the cost of a blip is a stale status until the next recovery or reconcile pass picks the row up. A retry policy tightens that window for live processes. It is a refinement, not a gap.

Why there is no transient retry today

EF's EnableRetryOnFailure was never a usable shortcut here, for two reasons:

  1. Several write paths open explicit transactions, which EF's retrying execution strategy rejects outright.
  2. More fundamentally, its classifier treats a command timeout as transient and replays the whole operation, and a timeout can arrive after the server already committed. Blind replay is unsafe in this codebase:
  • The CAS advance does not consume its own token. usp_UpdateCurrentRunCas guards on ScheduleVersion = @ExpectedScheduleVersion, then increments CurrentRunCount without changing that version (the Postgres CTE has the same shape). After an ambiguous commit the identical retry passes the guard again: one run counted twice, MaxRuns spent early, a series can end before its time.
  • SetStatus at audit Full is not idempotent: each execution inserts a fresh StatusAudit row and re-stamps LastExecutionUtc.
  • A PK violation on a Persist retry proves prior success only when it is exactly PK_QueuedTasks. A TaskKey collision or the (ParentTaskId, ScheduledExecutionUtc) occurrence index firing is a different event and must not be read as "already inserted".

The policy to build (per operation, never a global flag)

  • Failure before the command was sent: bounded retry, always safe.
  • Persist ambiguous: read back by PK, verify the row's identity, then decide.
  • Status write ambiguous: read the row; retry only if the desired state is not already there.
  • CAS advance: automatic retry only if the operation is changed to consume its token (bump ScheduleVersion in the same statement); until then, verify-then-retry.
  • A caller-requested cancellation never turns into a retry.
  • Outage behavior: few attempts, exponential backoff with jitter, and a global budget or circuit breaker so a recovering database is not flooded by its own clients.

Gate before enabling anything

Fault injection before, during and after commit on every write shape, plus a sustained-outage run proving the budget holds and recovery still converges. EF's own guidance treats a commit failure as unknown state and recommends verification over replay; that is the bar.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority: lowMeasured and understood, but nobody is currently affected

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions