Outbox
Grove writes domain events to a transactional outbox before they are published to downstream consumers. This gives you exactly-once semantics with respect to the originating action: either the aggregate state change and the outbox row commit together, or neither does. A background poller then dispatches outbox rows to consumers.
Prerequisites: Workflows Overview, familiarity with records and events.
What you'll learn: The grove_outbox schema, the claim-lease protocol, the idempotency contract, and how to reason about delivery semantics.
Why an outbox?
When a Grove action emits an event, two things need to happen atomically:
- The aggregate state change and event log entry are written to the database.
- Downstream consumers (triggers, webhooks, external services) find out about the event.
Doing (2) inside the same transaction as (1) would couple correctness to the availability of every downstream consumer; doing it outside the transaction risks losing events when a process crashes between the commit and the dispatch. The outbox pattern solves this by making (2) a database write inside the transaction — so it commits atomically with (1) — and then relying on an independent poller to deliver the written row to consumers.
In Grove, outbox writes and state updates share the same transaction. Once the action commits, the event is durably queued even if the process dies before dispatch.
The grove_outbox table
Outbox rows live in a single shared system table, grove_outbox (the grove_ prefix is configurable; see Database Backends):
| Column | Type | Purpose |
|---|---|---|
pk | auto-increment | Physical ordering; sort key for dispatch |
project_id | TEXT | Project the event belongs to; scopes the poller's claim query |
resource | TEXT | Module/resource that produced the event |
aggregate_id | TEXT | Aggregate the event is about |
event_name | TEXT | Event type (e.g. "created") |
schema_version | INTEGER | Event schema version |
dedup_key | TEXT | Caller-supplied idempotency key (empty string by default) |
payload_json | TEXT | Event payload |
meta_json | TEXT | Event metadata (who, when, correlation IDs) |
claim_id | TEXT? | Set when a poller claims the row; NULL otherwise |
claimed_until | BIGINT? | Epoch seconds; the claim expires at this time |
created_at | BIGINT | Epoch milliseconds when the row was written |
dispatched_at | BIGINT? | Epoch milliseconds when the poller successfully delivered the row; NULL for undelivered rows |
Two invariants are enforced by the schema:
- Unique on
(project_id, dedup_key). If a caller writes two outbox rows for the same project with the same non-emptydedup_key, the second write fails. This gives you idempotency: re-running a command that supplied adedup_keyproduces at most one outbox row. - Index on
(project_id, dispatched_at, claimed_until, pk). The poller's claim query uses this index to find undispatched, unclaimed rows inpkorder efficiently.
The claim-lease dispatch protocol
Outbox rows are delivered by a background poller — ClaimingOutboxPoller in grove-server. Multiple poller instances can run concurrently across replicas without duplicating work, because each claim is leased.
Claiming a batch
The poller atomically marks a batch of undispatched rows with a claim identifier and a deadline:
UPDATE grove_outbox
SET claim_id = :my_claim_id,
claimed_until = :now + 60
WHERE project_id = :project
AND dispatched_at IS NULL
AND (claimed_until IS NULL OR claimed_until < :now)
LIMIT :batch_size
RETURNING pk, ..., payload_json, meta_json;
Any row whose claim_id is set is effectively "locked" to that poller for the next 60 seconds. Any other poller looking for work sees claimed_until >= now and skips that row.
Delivering the batch
For each claimed row, the poller dispatches to the matching consumer (trigger, webhook, external service). On success, it marks the row dispatched:
UPDATE grove_outbox
SET dispatched_at = :now_ms
WHERE pk = :row_pk
AND claim_id = :my_claim_id;
The claim_id check ensures a zombie poller whose lease expired cannot mark a row dispatched after another poller has reclaimed it.
Lease expiration
If the poller dies or is partitioned before dispatching, its lease expires after 60 seconds and other pollers become eligible to claim the row again. This is the only way a row ever gets retried: Grove does not attempt in-process retries; a failed dispatch is a dropped claim, and the row becomes re-claimable on the next poll.
Delivery semantics
The outbox gives you the following guarantees:
| Guarantee | Details |
|---|---|
| At-least-once delivery | A successfully committed outbox row will eventually be dispatched. Dispatch can repeat if a lease expires before the row is marked dispatched. |
| Per-aggregate ordering | Because pk is monotonic and claims process in pk order, events from a single aggregate are delivered in the order they were committed. Across aggregates, ordering is per-project but not globally deterministic. |
| Transactional coupling to state | An outbox row exists if and only if the matching aggregate change committed. There is no "event without a state change" or "state change without an event" window. |
Consequences for consumers
- Expect duplicates. A consumer that cannot tolerate duplicates must dedupe by
(project_id, resource, aggregate_id, event_name, schema_version)or bydedup_keywhen provided. - Design handlers to be idempotent. The most common mistake is a trigger that creates a row for each event without checking; a re-delivery then creates a duplicate row.
- Respect ordering within an aggregate. If your consumer tracks "last processed position," track it per
(resource, aggregate_id), not globally.
Supplying dedup_key
When a command is expected to be retried by the caller — for example, a webhook receiver that may re-deliver the same payload — pass a stable dedup_key when invoking the action. If the action runs twice with the same key, the second outbox write fails on the unique constraint, the command is rolled back, and no duplicate event is emitted.
The mechanics of supplying dedup_key through the action/command layer are currently part of the runtime's inbound request envelope; see the HTTP route documentation for where to set it on a route handler.
Operational tuning
The outbox poller's defaults (batch size, poll interval, lease duration) are chosen to give good throughput without hammering the database. For high-volume projects or latency-sensitive consumers, the poller accepts tuning at construction time:
| Setting | Default | Trade-off |
|---|---|---|
| Lease duration | 60 seconds | Shorter = faster recovery from dead pollers; longer = more tolerance for slow dispatches |
| Batch size | 100 | Larger = fewer DB round-trips; smaller = fairer share across project_ids when many are competing |
| Poll interval | 100 ms | Lower = lower dispatch latency; higher = less DB load |
These settings affect dispatch latency, not correctness. A row will be delivered at least once regardless.
Debugging
Common outbox questions and how to answer them:
| Question | Query |
|---|---|
| How many events are queued but not yet delivered? | SELECT count(*) FROM grove_outbox WHERE dispatched_at IS NULL AND project_id = '<your-project>' |
| Is a specific poller stuck? | SELECT claim_id, count(*), min(claimed_until) FROM grove_outbox WHERE dispatched_at IS NULL AND claimed_until > <now_unix> GROUP BY claim_id |
| Are there events stuck behind a lease that never expires? | SELECT pk, claim_id, claimed_until FROM grove_outbox WHERE dispatched_at IS NULL AND claimed_until < <now_unix> - 300 LIMIT 20 — rows with expired leases that haven't been reclaimed |
See Also
- Database Backends — where the outbox table lives alongside per-resource state
- Triggers — outbox rows feed event triggers
- Activities — activities can observe their own outbox-backed invocations
- Workflows Overview — how outbox delivery plugs into workflow orchestration