Skip to content

Materialize a window in one commit instead of one round trip per occurrence #49

Description

@GiampaoloGabba

Rescoped after measurement. The original issue proposed two changes; one of them was measured and
rejected, so what remains is the bulk write. Numbers, method and reproduction steps are in the
measurement comment below.

The problem

A materialization pass writes one occurrence per round trip: MaterializeOccurrence is called once per
granted slot, each with its own command against the store (OccurrenceMaterializer.cs, the grant loop).
A schedule that materializes a real window pays that once per child, sequentially.

What a batched write is worth

OMB compares a 500-slot window written one round trip at a time against the same window in one commit,
on Postgres, same table shape, same unique index, same FOR UPDATE on the schedule row, same
compare-and-swap on version and cursor. The bulk arm is the provider's own writable CTE with the single
row replaced by an unnest of arrays.

Arm Throughput Service time per occurrence
single, today's path 7,681/s 2,014 µs
bulk, one commit per window 85,274/s 177 µs

11x the throughput, 91% off the per-occurrence service time.

That is a conflict-free happy-path upper bound, not a gain an implementation would reach:

  • The bulk arm returns one count for the window. A real member owes a per-slot outcome, the per-slot
    metadata and the audit rows.
  • On a partial conflict it advances the cursor past the whole window and only then learns that fewer
    rows went in. A real member cannot do that.
  • The single arm goes through EF (a pooled context and command per occurrence) while the bulk arm holds
    one raw connection per lane, so part of the gap is per-call storage overhead rather than round trips.

Before implementing, re-measure the single arm over the same raw connection the bulk arm uses. That
separates the batching win from the EF overhead, and only the first is what this issue is about.

Who this helps

Only schedules that have opened their window. MaxPendingOccurrences defaults to 1, meaning strictly
serial (MisfireSettings.cs:36), and capacity is MaxPendingOccurrences - activeOccurrences
(DueSlotEnumerator.cs:111). With the default, a pass grants one occurrence, there is nothing to batch,
and this change is worth nothing. It pays for applications that deliberately raised that cap, and it
scales with how far they raised it.

What was measured and rejected

The other half of the original proposal was to prepare an immutable template once per pass, so that the
DI scope, the handler resolution and the payload serialization are not repeated per child. Measured, the
whole preparable part is 631 ns with no payload and 818 ns at 1 KB, against 793.5 µs per occurrence on
SQLite and 2,089 µs on Postgres: about 0.1%, and under 1% even with a 64 KB payload, where the cost
is writing the bytes rather than serializing them. With the default window it saves nothing at all,
since one grant per pass has nothing to amortize. An end-to-end subtraction could not resolve it above
noise at either concurrency.

That half is not worth writing. This issue does not cover it.

Design constraints for whoever picks it up

  • Per-slot compare-and-swap semantics have to survive: the unique index on (schedule, slot) is what
    decides duplicates today, and a batched write still has to report which individual slot lost a race.
  • ITaskStorage is public and already carries capability flags with default implementations
    (SupportsDurableOccurrences, SupportsScheduleVersioning). A bulk member should follow that pattern
    so custom storage implementations keep compiling and fall back to the per-slot path.
  • Each provider writes at its own tier today (procedures on SQL Server and MySQL, a writable CTE on
    Postgres, one SaveChanges on the EF base). A bulk member should keep that shape rather than force a
    common denominator.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance / allocation optimizationspriority: lowMeasured and understood, but nobody is currently affected

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions