Add large-payload blob auto-purge (opt-in singleton job, worker/SDK side) - #758
Add large-payload blob auto-purge (opt-in singleton job, worker/SDK side)#758YunchuWang wants to merge 21 commits into
Conversation
82ae04d to
0ac2dc3
Compare
0752610 to
cab0e9a
Compare
cab0e9a to
c05b15a
Compare
Large orchestration payloads are externalized to Azure Blob Storage as `blob:v1:<container>:<blobName>` tokens. The DTS backend stores those tokens but cannot delete the backing blobs (it has no storage credentials) — only this SDK can. This adds an opt-in, whole-scheduler singleton durable entity + orchestration job (mirroring src/ExportHistory) that drains payload rows the backend has soft-deleted and deletes their blobs, then acks so the backend can hard-delete the rows. Design: - PayloadStore.DeleteAsync is virtual (default throws NotSupportedException so it is non-breaking for existing external subclasses); BlobPayloadStore overrides it to decode the token and call DeleteIfExistsAsync (idempotent). - BlobPurgeJob (TaskEntity singleton): Create is a no-op when already Active so racing client processes don't disturb the running job; Run starts a fixed-id orchestrator. - BlobPurgeJobOrchestrator (perpetual): fetch a batch of tombstones, delete the blobs with capped parallelism, ack the successful deletions (failed tokens stay tombstoned to retry), idle on a timer when empty, ContinueAsNew periodically. - ExecuteBlobPurgeJobOperationOrchestrator bridges client -> entity. - Two new unary RPCs on TaskHubSidecarService: GetTombstonedPayloads / AckPurgedPayloads (authoritative proto follow-up: microsoft/durabletask-protobuf#76). - LargePayloadStorageOptions gains AutoPurge (opt-in, default false) and PayloadPurgeBatchSize (default 500). - Client-side BlobPurgeJobStarter (IHostedService) ensures the singleton job when AutoPurge is enabled, without blocking host startup. Worker always registers the entity/orchestrators/activities so a client-enabled job has something to run. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
c05b15a to
306d19f
Compare
…er simplification - Drop the `Dto` suffix now that the payload records are first-class public types in `Microsoft.DurableTask.Client` (`TombstonedPayload`, `PayloadPurgeAck`). - Collapse the magic `500` batch-size literal into a single `BlobPurgeConstants.DefaultBatchSize` used everywhere. - Rename `BlobPurgeJobStatus.Stopped` -> `Pending` (still the zero value) and remove the dead `Failed` member (nothing ever set it; the job self-heals). - Make the perpetual orchestrator self-heal: wrap each cycle in try/catch so a transient backend/entity/activity failure logs, backs off, and continues instead of failing the orchestration and killing the eternal loop. - Ack poison tokens: `DeleteExternalBlobActivity` now returns a three-way `BlobDeleteResult` (Deleted/Discarded/Retry). Malformed tokens are discarded and acked so the backend can clear the stuck row instead of re-streaming it forever; transient failures stay tombstoned to retry. - Replace the single-value `BlobPurgeJobCreationOptions` record with a plain `int` on `BlobPurgeJob.Create`. - Guard the client fetch RPC: `GetTombstonedPayloadsAsync` throws `ArgumentOutOfRangeException` unless `0 < limit < 1000`. - Simplify `BlobPurgeJobStarter` to a fixed-instance-id fire-once: drop the entity-active pre-check and schedule the Create bridge once with a fixed instance id, retrying only until the backend is reachable. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 28 out of 28 changed files in this pull request and generated no new comments.
Suppressed comments (2)
src/Extensions/AzureBlobPayloads/AutoPurge/Orchestrations/BlobPurgeJobOrchestrator.cs:149
DeleteExternalBlobActivityis invoked without thePurgeActivityRetryPolicy. The other purge activities use this retry policy and the PR description says activities use a small retry policy; without it, infra/worker-level activity failures will bubble up to the outer loop and delay a whole cycle instead of retrying the delete call in-place.
BlobDeleteResult result = await context.CallActivityAsync<BlobDeleteResult>(
nameof(DeleteExternalBlobActivity),
tombstone.Token);
src/Extensions/AzureBlobPayloads/DependencyInjection/DurableTaskClientBuilderExtensions.AzureBlobPayloads.cs:47
- This method calls the user-provided
configuredelegate twice (once viaServices.Configure(...)and once directly viaconfigure(probe)) to decide whether to register the hosted auto-purge starter. If the delegate has side effects (reads config with caching, increments counters, etc.) this can produce unexpected behavior. Consider avoiding the second invocation by always registering the starter and having it no-op whenAutoPurgeis false (or by introducing an explicit opt-in API for auto-purge).
// Conditional DI: register the auto-purge starter only when the caller opted into auto-purge. Peek the
// flag now by running the configure delegate against a probe (options configurators are pure setters).
LargePayloadStorageOptions probe = new();
configure(probe);
if (probe.AutoPurge)
{
RegisterBlobPurgeJobStarter(builder);
}
…ames with review spec
Rename the cross-account orphan log to BlobPurgeDeleteUnreachable (EventId 821,
Error) with the reviewer-specified message, mirror that exact wording in the
DeleteExternalBlobActivity PayloadStorageException catch, reword the
BlobPayloadStore cross-account message ("cannot delete in another account"), and
rename the three v2 DeleteAsync tests to the reviewer-specified names. No
behavior change.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 44c27836-c49c-45fd-ae2b-3309f9c3f0f0
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 28 out of 28 changed files in this pull request and generated 1 comment.
Suppressed comments (2)
src/Extensions/AzureBlobPayloads/DependencyInjection/DurableTaskClientBuilderExtensions.AzureBlobPayloads.cs:44
- The
configuredelegate is executed immediately on a probe instance (configure(probe)) to decide whether to register the hosted purge starter. This changes the normal Options/DI semantics (the delegate will now run twice: once here and once during options resolution) and can cause unexpected side effects or early exceptions if callers do anything beyond pure property setters (e.g., read environment/config, throw on missing settings, mutate shared state). Consider registering the hosted service unconditionally and having it no-op whenAutoPurgeis false (or using another runtime check) so the caller-providedconfiguredelegate is only executed by the options pipeline.
// Conditional DI: register the auto-purge starter only when the caller opted into auto-purge. Peek the
// flag now by running the configure delegate against a probe (options configurators are pure setters).
LargePayloadStorageOptions probe = new();
configure(probe);
if (probe.AutoPurge)
src/Extensions/AzureBlobPayloads/AutoPurge/Client/BlobPurgeJobStarter.cs:69
BlobPurgeJobStarterallocates aCancellationTokenSourcebut never disposes it. Disposing the CTS on stop avoids holding onto timer registrations and aligns with typicalIHostedServicecleanup patterns.
public async Task StopAsync(CancellationToken cancellationToken)
{
this.cts?.Cancel();
Task? pending = this.ensureTask;
| catch (PayloadStorageException ex) | ||
| { | ||
| // The token is well-formed but points at a storage account this worker's credential cannot reach | ||
| // (cross-account without AAD). Retrying can never succeed from this process, and because the backend | ||
| // streams tombstones with an uncursored TOP(N) query, leaving it un-acked would re-serve the same | ||
| // token every cycle and permanently block the purge pipeline. Discard it so the row is cleared, and | ||
| // log at Error so an operator can reclaim the orphaned blob out-of-band. | ||
| this.logger.BlobPurgeDeleteUnreachable(ex, input); | ||
| return BlobDeleteResult.Discarded; | ||
| } | ||
| catch (Exception ex) when (ex is not OutOfMemoryException and not StackOverflowException) | ||
| { | ||
| // Transient failure: leave the payload tombstoned so a later purge cycle can retry it. A single | ||
| // bad token must not fail the whole batch. | ||
| this.logger.BlobPurgeDeleteFailed(ex, input); | ||
| return BlobDeleteResult.Retry; | ||
| } |
Decision 1: the auto-purge job declines to act on legacy v1 tokens. DeleteExternalBlobActivity.RunAsync discards v1 tokens up front - logged at error as the operator recovery pointer - without calling the store, because a v1 token identifies no storage account and a delete cannot be verified. BlobPayloadStore.TokenPrefixV1 is made internal so the activity can reference it; BlobPayloadStore.DeleteAsync is unchanged and still honors v1 for direct callers. Decision 2: do not start the purge job when the registered store cannot delete. BlobPurgeJobStarter now takes the resolved PayloadStore and no-ops at startup (logged at error, does not throw) when it is not a BlobPayloadStore. Defense in depth: the activity catches NotSupportedException and returns Retry to keep the payload tombstoned rather than acking a row whose blob was never deleted. Adds Logs 822/823/824, XML doc updates, a per-task-hub singleton phrasing fix, and focused tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 44c27836-c49c-45fd-ae2b-3309f9c3f0f0
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (3)
src/Extensions/AzureBlobPayloads/DependencyInjection/DurableTaskClientBuilderExtensions.AzureBlobPayloads.cs:44
- UseExternalizedPayloads invokes the user-provided configure delegate twice (once via Services.Configure and again on a probe instance) to detect AutoPurge. If configure has side-effects (e.g., reads env vars, performs validation, logs, throws on missing settings), those side-effects will happen twice and can be surprising.
// Conditional DI: register the auto-purge starter only when the caller opted into auto-purge. Peek the
// flag now by running the configure delegate against a probe (options configurators are pure setters).
LargePayloadStorageOptions probe = new();
configure(probe);
if (probe.AutoPurge)
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:28
- BlobPurgeJob.Create stores purgeBatchSize without validating it. If an invalid value (e.g., 0 or >1000) is ever provided, the orchestrator will repeatedly fail when calling GetTombstonedPayloadsAsync (GrpcDurableTaskClient enforces 1..1000), leaving the singleton job stuck in an Active-but-broken loop.
public void Create(TaskEntityContext context, int purgeBatchSize)
{
if (this.State.Status == BlobPurgeJobStatus.Active)
{
logger.BlobPurgeJobAlreadyRunning(context.Id.Key);
src/Extensions/AzureBlobPayloads/AutoPurge/Orchestrations/BlobPurgeJobOrchestrator.cs:148
- DeleteExternalBlobActivity is invoked without the purge retry policy (unlike GetTombstonedPayloadsActivity/AckPurgedPayloadsActivity). This contradicts the stated design of using a small retry policy for purge activities and makes transient activity failures fail the cycle immediately instead of being retried per PurgeActivityRetryPolicy.
BlobDeleteResult result = await context.CallActivityAsync<BlobDeleteResult>(
nameof(DeleteExternalBlobActivity),
tombstone.Token);
Address three reviewer findings on the blob payload auto-purge starter: - Finding 1: register BlobPurgeJobStarter unconditionally from UseExternalizedPayloadsCore (both overloads) instead of probing the configure delegate at registration time. The probe double-invoked user code and missed every enable path except the inline delegate (services.Configure, config binding, PostConfigure, parameterless overload). The AutoPurge decision now happens in StartAsync, once options are fully resolved, and no-ops silently when disabled. - Finding 2: inject IDurableTaskClientProvider and resolve the client by builder name lazily in StartAsync (after both gates) so named clients get their own client and no DurableTaskClient is constructed at host start for apps that externalize payloads without auto-purge. - Finding 3: pass the purge retry policy to the delete activity call in BlobPurgeJobOrchestrator.DeleteOneAsync, matching the other two calls. Tests: update starter tests for the new ctor, add a disabled-path test asserting no client resolution or logging, and add DI regression tests covering the services.Configure enable path and single configure invocation. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 44c27836-c49c-45fd-ae2b-3309f9c3f0f0
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (1)
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:36
purgeBatchSizeis stored without validating range. If the entity is created with 0/negative (e.g., via a direct entity call), the orchestrator will pass that value toGetTombstonedPayloadsAsync, which throws forlimit <= 0, causing the purge job to spin/fail every cycle. Consider validating here to prevent creating a permanently-broken job state.
this.State.Status = BlobPurgeJobStatus.Active;
this.State.PurgeBatchSize = purgeBatchSize;
this.State.CreatedAt ??= DateTimeOffset.UtcNow;
this.State.LastModifiedAt = DateTimeOffset.UtcNow;
this.State.LastError = null;
Item A: add the SA1600 doc comment to the internal TokenPrefixV1 const in BlobPayloadStore.cs, restoring the multi-TFM warning baseline to 201. Item B: add an else branch to BlobPurgeJobOrchestrator.RunAsync for the case where a batch produced no acks. DeleteExternalBlobActivity reports storage failures as a BlobDeleteResult.Retry return value rather than an exception, so the activity retry policy never engages and there is no backoff on that path. Combined with the backend's uncursored TOP(N) tombstone query, an immediate continue would refetch the identical rows and re-attempt the identical deletes in a tight loop for the duration of a storage outage. Back off on ErrorBackoff (the existing one-minute timer) before the next cycle. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 44c27836-c49c-45fd-ae2b-3309f9c3f0f0
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (2)
src/Extensions/AzureBlobPayloads/AutoPurge/Client/BlobPurgeJobStarter.cs:95
- StopAsync cancels the CTS but never awaits/observes the background ensure task (it only awaits Task.WhenAny) and never disposes the CTS. This can leave exceptions unobserved and can allow the ensure loop to keep running past shutdown if it doesn’t complete promptly. Prefer awaiting the task with a bounded cancellation token and dispose/clear the CTS/task (pattern matches other hosted services in the repo, e.g. SandboxActivityWorkerRegistrationHostedService.StopAsync).
public async Task StopAsync(CancellationToken cancellationToken)
{
this.cts?.Cancel();
Task? pending = this.ensureTask;
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:65
- BlobPurgeJob.Run schedules a fixed-instance-id orchestrator without any exception handling. In this repo, entities that start orchestrations wrap ScheduleNewOrchestration in try/catch to avoid rolling back the entity operation (and losing the start action) on transient/expected errors (see src/ExportHistory/Entity/ExportJob.cs:259-293 and src/ScheduledTasks/Entity/Schedule.cs:269-306). Consider aligning with that pattern and recording LastError when scheduling fails.
string instanceId = BlobPurgeConstants.GetOrchestratorInstanceId(context.Id.Key);
StartOrchestrationOptions startOrchestrationOptions = new(instanceId);
context.ScheduleNewOrchestration(
new TaskName(nameof(BlobPurgeJobOrchestrator)),
The DTS backend now hard-deletes legacy blob:v1: payload rows instead of tombstoning them, so the SDK will normally never be handed a v1 token. Update the prose that described v1-tombstoning as the expected path. - LargePayloadStorageOptions.AutoPurge <remarks> (public API): v1-backed blobs are not reclaimed and remain in storage exactly as before auto-purge existed (not a new leak); the backend removes their rows normally. - Logs.cs EventId 822: reworded to read as unexpected (older backend, or a row tombstoned before the fix); keeps Error level, EventId 822, and the full token. - DeleteExternalBlobActivity <remarks> item + inline gate comment: note the v1 branch is now a defensive guard for backend version skew, not the expected path. Prose only; no behavioural change. The only non-comment edit is the 822 message string. Multi-TFM warning baseline unchanged (201); AzureBlobPayloads tests 42/42. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 44c27836-c49c-45fd-ae2b-3309f9c3f0f0
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (4)
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:36
- BlobPurgeJob.Create persists purgeBatchSize without validating it. If a caller passes 0 or > MaxBatchSize (e.g., via the public ExecuteBlobPurgeJobOperationOrchestrator bridge), the purge orchestrator will repeatedly fail because GrpcDurableTaskClient.GetTombstonedPayloadsAsync throws for limit <= 0, resulting in an endless backoff loop with no progress. Consider validating the range here (at least for non-Active creates) so invalid values fail fast instead of permanently wedging the singleton job.
this.State.Status = BlobPurgeJobStatus.Active;
this.State.PurgeBatchSize = purgeBatchSize;
this.State.CreatedAt ??= DateTimeOffset.UtcNow;
this.State.LastModifiedAt = DateTimeOffset.UtcNow;
this.State.LastError = null;
src/Client/Grpc/GrpcDurableTaskClient.cs:690
- AckPurgedPayloadsAsync also doesn't translate StatusCode.Unimplemented. If called against an older sidecar, it will throw RpcException instead of a consistent CLR exception (other optional RPCs in this client translate Unimplemented). Add an Unimplemented catch here too.
catch (RpcException e) when (e.StatusCode == StatusCode.Cancelled)
{
throw new OperationCanceledException(
$"The {nameof(this.AckPurgedPayloadsAsync)} operation was canceled.", e, cancellation);
}
test/Extensions/AzureBlobPayloads.Tests/AutoPurge/BlobPurgeJobTests.cs:78
- This test locks in accepting purge batch size = 0, but 0 will cause GetTombstonedPayloadsAsync(limit) to throw (GrpcDurableTaskClient enforces 1..1000). If BlobPurgeJob.Create is hardened to validate its input (recommended), update this test to assert the invalid batch size is rejected.
public async Task Create_StoresBatchSizeVerbatim_WithoutCoercion()
{
// Arrange - the batch size is validated once at specification (LargePayloadStorageOptions), so the
// entity trusts its input and performs no coercion of its own. A zero here is stored as-is, proving
// the previous non-positive-to-default fallback was removed.
src/Client/Grpc/GrpcDurableTaskClient.cs:648
- GetTombstonedPayloadsAsync currently only translates StatusCode.Cancelled. If the sidecar/backend is older and doesn't implement the new RPC, gRPC will return StatusCode.Unimplemented and this method will leak a raw RpcException (inconsistent with other methods in this type that translate Unimplemented to a CLR exception). Add a StatusCode.Unimplemented translation so callers get a predictable exception type.
This issue also appears on line 686 of the same file.
catch (RpcException e) when (e.StatusCode == StatusCode.Cancelled)
{
throw new OperationCanceledException(
$"The {nameof(this.GetTombstonedPayloadsAsync)} operation was canceled.", e, cancellation);
}
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated 1 comment.
Suppressed comments (2)
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:35
- BlobPurgeJob.Create persists purgeBatchSize without validating the range. The rest of the pipeline assumes 1..1000 (LargePayloadStorageOptions enforces this and GrpcDurableTaskClient.GetTombstonedPayloadsAsync throws for <=0 or >1000). Because the entity is driven via a public bridge orchestrator (ExecuteBlobPurgeJobOperationOrchestrator + BlobPurgeJobOperationRequest), external callers can still invoke Create with an invalid value, leading to a perpetual purge job that just backs off on errors. Consider validating purgeBatchSize here (or coercing to a safe default) to keep the job robust against malformed inputs.
public void Create(TaskEntityContext context, int purgeBatchSize)
{
if (this.State.Status == BlobPurgeJobStatus.Active)
{
logger.BlobPurgeJobAlreadyRunning(context.Id.Key);
return;
}
this.State.Status = BlobPurgeJobStatus.Active;
this.State.PurgeBatchSize = purgeBatchSize;
this.State.CreatedAt ??= DateTimeOffset.UtcNow;
this.State.LastModifiedAt = DateTimeOffset.UtcNow;
src/Extensions/AzureBlobPayloads/AutoPurge/Client/BlobPurgeJobStarter.cs:74
- The auto-purge startup gate assumes only BlobPayloadStore can delete payloads ("if (this.store is not BlobPayloadStore)"). However PayloadStore.DeleteAsync is virtual and external implementations can override it to support deletion (including blob-token-compatible implementations). As written, such stores will be incorrectly treated as non-deletable and AutoPurge will never start. Consider checking delete capability (override) instead of concrete type.
// Auto-purge deletes blobs through the store, but PayloadStore.DeleteAsync is virtual and its base
// implementation throws NotSupportedException. A store that cannot delete would fail every single
// payload, so refuse to start the job rather than spin against the backend - and rather than ack rows
// whose blobs were never deleted, which would destroy the backend's record of what still needs cleanup.
// This is a configuration error and is surfaced at startup, where it is cheapest to notice.
if (this.store is not BlobPayloadStore)
{
this.logger.BlobPurgeStoreCannotDelete(this.store.GetType().FullName);
return Task.CompletedTask;
}
| static void RegisterBlobPurgeJobStarter(IDurableTaskClientBuilder builder) | ||
| { | ||
| string builderName = builder.Name; | ||
| builder.Services.AddSingleton<IHostedService>(sp => new BlobPurgeJobStarter( | ||
| sp.GetRequiredService<IDurableTaskClientProvider>(), | ||
| sp.GetRequiredService<PayloadStore>(), | ||
| sp.GetRequiredService<IOptionsMonitor<LargePayloadStorageOptions>>(), | ||
| builderName, | ||
| sp.GetRequiredService<ILogger<BlobPurgeJobStarter>>())); | ||
| } |
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 29 out of 29 changed files in this pull request and generated no new comments.
Suppressed comments (2)
src/Extensions/AzureBlobPayloads/AutoPurge/Entity/BlobPurgeJob.cs:36
purgeBatchSizeis stored verbatim into entity state without validation. If a caller invokesBlobPurgeJob.Createwith 0/negative or >1000, the orchestrator will later fail when calling GetTombstonedPayloads (the gRPC client enforces 1..1000), causing the job to back off and loop forever without making progress. Consider validating the range here (or coercing to a safe default) even if the starter path uses validated options, since the entity operation is callable independently.
this.State.Status = BlobPurgeJobStatus.Active;
this.State.PurgeBatchSize = purgeBatchSize;
this.State.CreatedAt ??= DateTimeOffset.UtcNow;
this.State.LastModifiedAt = DateTimeOffset.UtcNow;
this.State.LastError = null;
src/Extensions/AzureBlobPayloads/AutoPurge/Client/BlobPurgeJobStarter.cs:74
- The auto-purge capability gate is currently
this.store is not BlobPayloadStore, which assumes only BlobPayloadStore can delete payloads. SincePayloadStore.DeleteAsyncwas intentionally added as a virtual extensibility point, a custom PayloadStore could support deletion without being BlobPayloadStore. Consider gating on whetherDeleteAsyncis overridden (capability) rather than the concrete store type (implementation), so delete-capable custom stores can still opt in.
if (this.store is not BlobPayloadStore)
{
this.logger.BlobPurgeStoreCannotDelete(this.store.GetType().FullName);
return Task.CompletedTask;
}
Summary
Large orchestration payloads are externalized to Azure Blob Storage by the
AzureBlobPayloadsextension asblob:v1:<container>:<blobName>tokens. The DTS backend stores those tokens but cannot delete the backing blobs — it has no storage credentials; only this SDK can. This PR implements the worker/SDK side of large-payload blob auto-purge.Design (opt-in singleton durable entity + orchestration job)
Instead of an always-on background stream, this mirrors the existing
src/ExportHistoryfeature: an opt-in, whole-scheduler singleton durable entity + orchestration job that drains soft-deleted payload rows the backend exposes and deletes their blobs, then acks so the backend can hard-delete the rows.PayloadStore.DeleteAsync— added as avirtualmethod (default throwsNotSupportedException, so it is non-breaking for existing external subclasses).BlobPayloadStoreoverrides it to decode the token and callDeleteIfExistsAsync(idempotent — deleting a missing blob is a no-op).BlobPurgeJob(TaskEntitysingleton) —Createis a no-op when already Active so racing client processes don't disturb the running job (intentionally softer thanExportJob.Create, which throws).Runstarts a fixed-instance-id orchestrator.BlobPurgeJobOrchestrator(perpetual) — each cycle: fetch a batch of tombstones, delete the blobs with capped parallelism (32), ack only the successful deletions (failed tokens stay tombstoned to retry), idle on a 1-minute timer when there's nothing to purge, andContinueAsNewevery 5 cycles to keep history small. Activities use a small retry policy.ExecuteBlobPurgeJobOperationOrchestrator— client → entity bridge (mirrors export).GetTombstonedPayloadsActivity,DeleteExternalBlobActivity(returnsfalse+ logs on failure so one bad token can't fail the batch),AckPurgedPayloadsActivity.DurableTaskClient(GetTombstonedPayloadsAsync/AckPurgedPayloadsAsync), overridden inGrpcDurableTaskClient— mirroring how ExportHistory addedListInstanceIdsAsync/GetOrchestrationHistoryAsync. The purge activities injectDurableTaskClientdirectly; no dedicated gRPC client is needed. AzureManaged reusesGrpcDurableTaskClient, so it inherits these methods with no extra client.LargePayloadStorageOptions.AutoPurge(opt-in, defaultfalse) andPayloadPurgeBatchSize(default500).BlobPurgeJobStarter(IHostedService) ensures the singleton job whenAutoPurgeis enabled, on a background task that does not block host startup and retries until the backend is reachable. The worker always registers the entity/orchestrators/activities (not gated onAutoPurge) so a client-enabled job always has something to execute.gRPC contract
Two new unary RPCs added to
TaskHubSidecarServiceinsrc/Grpc/orchestrator_service.proto(worker is the client; wire paths/TaskHubSidecarService/GetTombstonedPayloadsand/AckPurgedPayloads):C# stubs are generated at build time by
Grpc.Tools(not committed). The authoritative proto change is a follow-up in microsoft/durabletask-protobuf#76; once it merges,src/Grpc/orchestrator_service.protoshould be re-synced from upstream (content identical to what's here). The vendored change lets this PR build/test standalone.Testing
dotnet build Microsoft.DurableTask.sln— succeeds, 0 errors.dotnet test test/AzureBlobPayloads.Tests— 11 passed (BlobPayloadStore delete/idempotency +BlobPurgeJob.Createno-op-when-Active + options defaults).dotnet test test/ExportHistory.Tests— 147 passed (shared patterns unaffected).dotnet test test/Client/Core.Tests+test/Client/Grpc.Tests— 43 + 47 passed (coreDurableTaskClient/GrpcDurableTaskClientedits).Notes / deviations
BlobPurgeJobStatus.Pendingis the0/default value (mirroringExportJobStatus.Pending=0) so a freshly initialized entity never appearsActive.RecordPurgedentity op to track a cumulativePurgedCount.Co-authored-by: Copilot App 223556219+Copilot@users.noreply.github.com