Outbox

Grove writes domain events to a transactional outbox before they are published to downstream consumers. This gives you exactly-once semantics with respect to the originating action: either the aggregate state change and the outbox row commit together, or neither does. A background poller then dispatches outbox rows to consumers.

Prerequisites: Workflows Overview, familiarity with records and events. What you'll learn: The grove_outbox schema, the claim-lease protocol, the idempotency contract, and how to reason about delivery semantics.


Why an outbox?

When a Grove action emits an event, two things need to happen atomically:

  1. The aggregate state change and event log entry are written to the database.
  2. Downstream consumers (triggers, webhooks, external services) find out about the event.

Doing (2) inside the same transaction as (1) would couple correctness to the availability of every downstream consumer; doing it outside the transaction risks losing events when a process crashes between the commit and the dispatch. The outbox pattern solves this by making (2) a database write inside the transaction — so it commits atomically with (1) — and then relying on an independent poller to deliver the written row to consumers.

In Grove, outbox writes and state updates share the same transaction. Once the action commits, the event is durably queued even if the process dies before dispatch.

The grove_outbox table

Outbox rows live in a single shared system table, grove_outbox (the grove_ prefix is configurable; see Database Backends):

ColumnTypePurpose
pkauto-incrementPhysical ordering; sort key for dispatch
project_idTEXTProject the event belongs to; scopes the poller's claim query
resourceTEXTModule/resource that produced the event
aggregate_idTEXTAggregate the event is about
event_nameTEXTEvent type (e.g. "created")
schema_versionINTEGEREvent schema version
dedup_keyTEXTCaller-supplied idempotency key (empty string by default)
payload_jsonTEXTEvent payload
meta_jsonTEXTEvent metadata (who, when, correlation IDs)
claim_idTEXT?Set when a poller claims the row; NULL otherwise
claimed_untilBIGINT?Epoch seconds; the claim expires at this time
created_atBIGINTEpoch milliseconds when the row was written
dispatched_atBIGINT?Epoch milliseconds when the poller successfully delivered the row; NULL for undelivered rows

Two invariants are enforced by the schema:

The claim-lease dispatch protocol

Outbox rows are delivered by a background poller — ClaimingOutboxPoller in grove-server. Multiple poller instances can run concurrently across replicas without duplicating work, because each claim is leased.

Claiming a batch

The poller atomically marks a batch of undispatched rows with a claim identifier and a deadline:

UPDATE grove_outbox
   SET claim_id = :my_claim_id,
       claimed_until = :now + 60
 WHERE project_id = :project
   AND dispatched_at IS NULL
   AND (claimed_until IS NULL OR claimed_until < :now)
 LIMIT :batch_size
RETURNING pk, ..., payload_json, meta_json;

Any row whose claim_id is set is effectively "locked" to that poller for the next 60 seconds. Any other poller looking for work sees claimed_until >= now and skips that row.

Delivering the batch

For each claimed row, the poller dispatches to the matching consumer (trigger, webhook, external service). On success, it marks the row dispatched:

UPDATE grove_outbox
   SET dispatched_at = :now_ms
 WHERE pk = :row_pk
   AND claim_id = :my_claim_id;

The claim_id check ensures a zombie poller whose lease expired cannot mark a row dispatched after another poller has reclaimed it.

Lease expiration

If the poller dies or is partitioned before dispatching, its lease expires after 60 seconds and other pollers become eligible to claim the row again. This is the only way a row ever gets retried: Grove does not attempt in-process retries; a failed dispatch is a dropped claim, and the row becomes re-claimable on the next poll.

Delivery semantics

The outbox gives you the following guarantees:

GuaranteeDetails
At-least-once deliveryA successfully committed outbox row will eventually be dispatched. Dispatch can repeat if a lease expires before the row is marked dispatched.
Per-aggregate orderingBecause pk is monotonic and claims process in pk order, events from a single aggregate are delivered in the order they were committed. Across aggregates, ordering is per-project but not globally deterministic.
Transactional coupling to stateAn outbox row exists if and only if the matching aggregate change committed. There is no "event without a state change" or "state change without an event" window.

Consequences for consumers

Supplying dedup_key

When a command is expected to be retried by the caller — for example, a webhook receiver that may re-deliver the same payload — pass a stable dedup_key when invoking the action. If the action runs twice with the same key, the second outbox write fails on the unique constraint, the command is rolled back, and no duplicate event is emitted.

The mechanics of supplying dedup_key through the action/command layer are currently part of the runtime's inbound request envelope; see the HTTP route documentation for where to set it on a route handler.

Operational tuning

The outbox poller's defaults (batch size, poll interval, lease duration) are chosen to give good throughput without hammering the database. For high-volume projects or latency-sensitive consumers, the poller accepts tuning at construction time:

SettingDefaultTrade-off
Lease duration60 secondsShorter = faster recovery from dead pollers; longer = more tolerance for slow dispatches
Batch size100Larger = fewer DB round-trips; smaller = fairer share across project_ids when many are competing
Poll interval100 msLower = lower dispatch latency; higher = less DB load

These settings affect dispatch latency, not correctness. A row will be delivered at least once regardless.

Debugging

Common outbox questions and how to answer them:

QuestionQuery
How many events are queued but not yet delivered?SELECT count(*) FROM grove_outbox WHERE dispatched_at IS NULL AND project_id = '<your-project>'
Is a specific poller stuck?SELECT claim_id, count(*), min(claimed_until) FROM grove_outbox WHERE dispatched_at IS NULL AND claimed_until > <now_unix> GROUP BY claim_id
Are there events stuck behind a lease that never expires?SELECT pk, claim_id, claimed_until FROM grove_outbox WHERE dispatched_at IS NULL AND claimed_until < <now_unix> - 300 LIMIT 20 — rows with expired leases that haven't been reclaimed

See Also