Exactly-Once Side Effects: What a Workflow Engine Can Actually Promise

A workflow step charges a card, the engine crashes before it saves the response, and on restart it has to decide whether retrying is safe. This is how my engine answers that with an eight-state receipt per effect, a deterministic receipt ID, and a conditional UPDATE in the database.

By Oleksii Vasylenko, Technical Lead · Published · Updated · 14 min read

Where this comes from. Orch8 is a durable workflow engine I built solo in Rust: every crate, every SDK, every integration. Ten-crate workspace, PostgreSQL and SQLite backends, server and on-device execution. The details below are the actual implementation, not a reference architecture.

A workflow step calls a payment API. The process dies after the HTTP request leaves and before the response is written anywhere. On restart, the engine finds a step that started and never finished. If it retries, it may charge the customer twice. If it does not retry, the order may sit unpaid forever. Every durable workflow engine has this window, and most of the marketing around "exactly-once" is about how it is described rather than how it is closed.

The claim that holds up is narrower. State transitions inside the engine's own database can be exactly-once, because one transaction owns them. Effects on somebody else's system can be at-most-once, because you can write down your intent before you send anything. And sometimes the outcome of a dispatched effect is unknown, in which case the only correct move is to stop and say so. Orch8 is built around that split. The guard that implements it opens with a doc comment saying it provides at-most-once safety rather than exactly-once semantics, and that once a dispatch is durably recorded, an ambiguous outcome blocks automatic retry until an operator or a verifier resolves the receipt.

The crash window between dispatching a side effect and recording its outcomeThe engine writes a receipt, sends a charge request, then dies before the response arrives. On restart the receipt reads Dispatched with an unknown outcome, retry is blocked, and only a reconciler or operator can settle it.Payment APIReceipt storeEngineprocess dies hereon restart the receipt saysDispatched, outcome unknownonly a reconciler or operatorcan settle this receiptwrite receipt (Prepared)1POST /charge2200 OK (nobody listening)3Dispatched → Unknown4retry blocked5
The request leaves, the process dies, the answer never arrives. From inside the engine, "the charge failed" and "the charge succeeded and the response was lost" look identical.

I am going to walk through what actually happens to one step in that window: which rows get written, in what order, what the SQL looks like, and what an operator sees afterwards. The general pattern transfers to any system that calls an external API from inside retryable work.

A quick model of the engine, because the rest of the article uses its vocabulary. A workflow definition is a sequence: a versioned JSON document containing blocks. The simplest block is a step, and a step names a handler such as http_request, llm_call, or a custom handler registered by the operator. A run of a sequence is an instance, stored as one row in task_instances with a state, a JSONB context, and a next_fire_at timestamp. The engine is a single Rust binary. Its scheduler runs a tick loop against Postgres (or SQLite for tests and on-device use): claim due instances, execute whatever blocks are ready, write the results, release.

Step execution lives in orch8-engine/src/handlers/step.rs. The function execute_step looks up the handler, resolves parameters, and then, before the handler runs, calls EffectGuard::begin with the tenant, instance, block ID, handler name, resolved params, and the attempt number. Whether that call returns a guard or None depends on the handler, which is the next section. When the handler returns, the same function calls guard.commit(&output) on success, guard.mark_unknown() on a timeout or an error, and only then writes the block output row that the scheduler uses to decide what to run next.

The attempt number matters more than it looks. It is derived in compute_attempt from the most recent block_outputs row for that block. A step that crashed mid-flight leaves an __in_progress__ sentinel as its latest row, and the engine resumes at the same attempt. A step whose handler returned a clean error leaves a __retry__ marker, and the next run is attempt plus one. The receipt logic relies on that distinction, so keep it in mind.

The usual representation of a side effect is a boolean, done or not done. A boolean has no way to say "we sent it and never heard back", and that is the case that produces incidents. So every side-effecting step gets a receipt row in effect_receipts, and the receipt moves through a state machine defined in orch8-types/src/continuity.rs:

pub enum EffectState {
    Planned,      // decided to do it
    Prepared,     // receipt durably written, not yet sent
    Dispatched,   // sent; outcome not yet known
    Committed,    // provider confirmed
    Unknown,      // sent, no answer
    Verified,     // reconciled against the provider afterwards
    Compensated,  // undone by a compensating action
    Abandoned,    // written off by an operator, with evidence
}

pub const fn can_transition_to(self, next: Self) -> bool {
    matches!(
        (self, next),
        (Self::Planned, Self::Prepared | Self::Abandoned)
            | (Self::Prepared, Self::Dispatched | Self::Abandoned)
            | (Self::Dispatched, Self::Committed | Self::Unknown)
            | (Self::Unknown,
               Self::Verified | Self::Committed | Self::Compensated | Self::Abandoned)
            | (Self::Committed | Self::Verified, Self::Compensated)
    )
}
orch8-types/src/continuity.rs. The enum and the transition table. Unknown can only leave through an operator or a verifier.
The eight states of an effect receiptPlanned, Prepared, Dispatched, then either Committed or Unknown. Unknown blocks automatic retry and resolves to Verified, Compensated, or Abandoned.receipt durablesentprovider confirmedcrash / no answerreconciledundonewritten offnew attemptPlannedPreparedDispatchedCommittedUnknownVerifiedCompensatedAbandonedblocks automatic retry
Planned, Prepared, Dispatched, then Committed or Unknown. Three of the eight states (Verified, Compensated, Abandoned) exist for people and reconciliation jobs rather than for the engine.

Unknown is the state the whole design is built around. A receipt arrives there in two ways. The handler returns an error or times out while the receipt is Dispatched, and the engine marks it Unknown because it cannot tell whether the provider acted. Or the process dies while the receipt is Dispatched, and the next EffectGuard::begin for the same block and attempt finds it and moves it to Unknown before refusing to proceed. Notice that Unknown has no transition back to Dispatched. The engine cannot re-send on its own. It can only wait for Verified, Committed, Compensated, or Abandoned, and each of those requires evidence from outside the engine.

Every transition is a compare-and-swap. advance calls receipt.transition(next), which rejects illegal moves, and then issues cas_effect_receipt with the state it expects to find. In Postgres that is one statement:

UPDATE effect_receipts
   SET state = $1, record = $2, updated_at = $3
 WHERE tenant_id = $4 AND id = $5 AND state = $6   -- $6 = expected state
orch8-storage/src/postgres/continuity.rs. The WHERE clause carries the expected state, so a stale caller changes zero rows and finds out.

If rows_affected is not one, some other process already moved the receipt, and the current caller is working from a stale view. The guard reloads the row and returns EngineError::EffectBlocked with the current state. A lost CAS is never treated as a signal to retry.

Here is the order of operations for a step with handler http_request and params { "url": "https://payments.example/charge", "idempotency_key": "order-7" }, following EffectGuard::begin in orch8-engine/src/effect_guard.rs.

  1. handler_has_side_effects("http_request") returns true, so a guard is built. For transform or log it would return false and the step would run with no receipt at all.
  2. ensure_effect_scope derives a ContinuityExecution for the instance. The continuity ID and runtime ID are SHA-256 digests of the instance ID under fixed domain strings, so two processes handling the same instance converge on one scope without coordinating. An instance that has never been enrolled in portable continuity still gets a ledger this way.
  3. find_unresolved_effect_receipt looks for a receipt on this continuity, instance, block, and attempt in state dispatched or unknown. If one exists, the crash case has happened. A dispatched receipt is moved to unknown and the call returns EffectBlocked. Nothing is sent.
  4. Otherwise deterministic_effect_id computes the receipt ID (next section), and a Planned receipt is built with the effect kind, a destination fingerprint, the caller's idempotency key if any, and a SHA-256 of the canonical request.
  5. ensure_effect_receipt upserts it: INSERT ... ON CONFLICT (id) DO UPDATE SET id = EXCLUDED.id RETURNING record. The no-op update is there so the statement always returns the row that won, whether it is ours or a concurrent process's. same_effect_identity then compares what came back with what we tried to insert. A mismatch means a hash collision or a bug, and the engine stops with InvalidConfig rather than reuse someone else's receipt.
  6. advance(Prepared) runs the CAS above. The receipt is now durable and the request has still not left the process.
  7. dispatch moves the receipt to Dispatched. With no at-most-once invariant declared this is another CAS. With one declared it becomes the conditional dispatch in the database, covered below.
  8. begin returns the guard. Only now does execute_step invoke the handler, and the HTTP request goes out.
  9. On success, commit extracts provider_receipt_id (it looks for provider_receipt_id, receipt_id, or request_id in the output), advances to Committed, and the step output is saved. On error or timeout, mark_unknown advances to Unknown and the error propagates.

Two writes before dispatch and one after, each a round trip to Postgres. That is the price of being able to answer "did it happen?" after a crash, and there is no version of this design where those writes are free. The receipt is written before dispatch because a receipt written after dispatch cannot record the case where the write itself never happens.

The evidence stored on the receipt is what a person will need later. The destination fingerprint hashes the handler name with whichever of url, provider, model, topic, queue, resource, and endpoint appear in the params. The request hash covers the whole canonicalised param object. When someone has to ask a payment provider "did you receive this?", those two fields plus the idempotency key and the provider receipt ID are the search terms.

A retry has to find the receipt the first attempt created, or the scheme collapses into one fresh receipt per attempt and no protection. A random UUID at dispatch time does that collapse. So the ID is a function of where the effect sits in the execution:

fn deterministic_effect_id(
    continuity_id: ContinuityId,   // which logical execution
    epoch: ExecutionEpoch,         // which ownership generation
    instance_id: InstanceId,       // which run
    block_id: &BlockId,            // which step
    attempt: u32,                  // which attempt
) -> EffectId {
    let mut hasher = Sha256::new();
    hasher.update(b"orch8-effect-v1\0");
    hasher.update(continuity_id.as_uuid().as_bytes());
    hasher.update(epoch.get().to_be_bytes());
    hasher.update(instance_id.into_uuid().as_bytes());
    hasher.update(block_id.as_str().as_bytes());
    hasher.update(attempt.to_be_bytes());
    // first 16 bytes of the digest, with the UUID v8 variant bits set
}
orch8-engine/src/effect_guard.rs. Same inputs, same receipt, from any process.

Including attempt in the derivation is a choice with a consequence I should spell out, because the first draft of this article got it slightly wrong. Recall that the attempt number comes from the latest block_outputs row. After a crash, the latest row is the __in_progress__ sentinel and the attempt does not advance, so the next begin computes the same ID, finds the dispatched receipt, and blocks. After a clean handler error, a __retry__ marker is written, the attempt advances, and the next begin computes a new ID. The old receipt sits in Unknown, the new one starts at Planned, and the retry is allowed to dispatch.

That is intentional. A provider that returned a 503 has told you nothing happened, and refusing to retry would turn every transient failure into a stuck workflow. But it also means the receipt ID alone gives you at-most-once per attempt, not per step. Cross-attempt protection comes from a separate mechanism, the at-most-once invariant, which compares request content rather than IDs. The two are complementary and neither is enough alone.

Wrapping every step in a receipt would add two round trips to work that cannot hurt anyone. A transform step rearranges JSON. A log step writes to the engine's own log. So the guard is only constructed for handlers that can reach the outside world, and the classification errs on the side of protecting too much:

const SIDE_EFFECT_BUILTINS: &[&str] = &[
    "http_request", "llm_call", "tool_call", "mcp_call", "agent",
    "emit_event", "send_signal", "self_modify",
    "memory_store", "memory_delete", "blob_put", "embed",
];

pub(crate) fn handler_has_side_effects(handler: &str) -> bool {
    SIDE_EFFECT_BUILTINS.contains(&handler)
        || !BUILTIN_HANDLER_NAMES.contains(&handler)   // not a builtin: assume yes
}
orch8-engine/src/release_diff.rs. The second clause is the one to notice.

A handler the engine does not recognise is a handler somebody wrote, and the engine cannot prove it does not send email. Defaulting to "protected" costs a little latency on custom steps. Defaulting to "safe" would exclude the code most likely to have effects the engine has never seen. The same list drives the semantic diff between two sequence versions: a change to a step using one of these handlers is classified as a side-effect risk rather than a behavioural change, and the release gate treats it accordingly.

The handler name also decides the receipt's kind. http_request is Http; llm_call, agent, and embed are Model; emit_event and send_signal are Message; the memory and blob handlers are Storage; and any non-builtin handler is Worker. Kinds matter because invariants are declared per kind.

A check in application code cannot stop two processes from dispatching the same effect. The check and the act are separated in time, and two engine nodes can both pass the check. When a sequence version declares an effect_at_most_once invariant with commit_guard: true for an effect kind, the transition to Dispatched stops being a plain CAS and becomes a transaction that decides which caller may send:

-- 1. Serialise callers for this (tenant, continuity, kind, destination, request)
SELECT pg_advisory_xact_lock(hashtextextended($1, 0));
--    $1 = 'effect|<tenant>|<continuity_id>|<kind>|<destination_fingerprint>|<request_sha256>'

-- 2. Read our receipt and every receipt with the same content
SELECT record FROM effect_receipts
 WHERE tenant_id = $1 AND continuity_id = $2
   AND (id = $3 OR (
         record->>'kind' = $4
     AND record->>'destination_fingerprint' = $5
     AND record->>'request_sha256' = $6))
 ORDER BY created_at, id
 FOR UPDATE;

-- 3. In Rust: if our receipt is missing or not Prepared        -> Stale
--    if another matching receipt is Dispatched/Committed/
--       Unknown/Verified, or is older than ours and not
--       Abandoned/Compensated                                   -> Duplicate

-- 4. Otherwise advance, still conditionally
UPDATE effect_receipts SET state = 'dispatched', record = $2, updated_at = $3
 WHERE tenant_id = $4 AND id = $5 AND state = 'prepared';
-- rows_affected != 1                                            -> Stale
orch8-storage/src/postgres/continuity.rs, dispatch_effect_receipt_at_most_once. Lock the guard key, read every competing receipt, decide, then update conditionally.

The advisory lock is the part I would not skip. Without it, two transactions can both read "no competing receipt", both update their own row, and both dispatch. With it, the second caller waits for the first to commit, then reads the first caller's dispatched receipt and returns Duplicate. There is a comment above that line about switching from hashtext to hashtextextended: the 32-bit hash collided often enough in tests that unrelated guard keys were serialised behind each other.

The "older than ours" clause in the duplicate check handles a race the lock alone does not. Two receipts for the same content can both exist in Prepared if two attempts raced through begin. Neither has dispatched. The rule is that the older receipt wins and the younger one reports Duplicate, so exactly one of them proceeds.

Three outcomes of a conditional dispatch at the storage boundaryDispatched means the caller owns the effect. Duplicate means an at-most-once invariant caught a second dispatch. Stale means the caller belongs to a superseded execution epoch.DispatchedDuplicateStaleconditional dispatchat the storage boundaryoutcomewe own itcall the providera second dispatch was attemptedrecord invariant violationcaller belongs to asuperseded epochblock and re-readCommitted or Unknown
Dispatched means this caller owns the effect. Duplicate means the invariant caught a second send of the same content. Stale means the caller belongs to a superseded epoch or lost the race.

Duplicate and Stale are kept apart on purpose. Duplicate produces EngineError::InvariantViolation with the invariant ID and block ID, and an invariant result row is appended so the failure shows up in the execution's evidence. Stale produces EffectBlocked with the current receipt state. A duplicate usually means a workflow bug (two steps charging the same card) or a retry after an Unknown. A stale dispatch usually means ownership of the execution moved to another runtime and an old owner tried to act. Folding those into one error would make two different incidents read the same in the logs.

Invariants are configuration on a sequence version, not global engine behaviour. load_effect_commit_guards reads the invariants for the instance's sequence and version and keeps only those with commit_guard: true and a rule matching the receipt's kind. Charging a card and appending to a log do not need the same protection, and the advisory lock plus FOR UPDATE read would be a tax on every step that does not need it.

EffectBlocked and InvariantViolation are not StepFailed errors, so the retry policy never sees them. In scheduler/step_exec.rs they fall through to the generic error arm: the scheduler writes an __error__ marker with retryable: false, transitions the instance from Running to Failed, fires the instance.failed webhook, and wakes the parent if this was a child instance. The workflow stops. Nothing is re-sent.

From there the path is manual, and docs/CONTINUITY_OPERATIONS.md spells it out. List the receipts with GET /instances/{id}/effects or the continuity-scoped equivalent. Match the effect ID, block ID, request hash, destination fingerprint, idempotency key, epoch, and attempt against the provider's records. If the provider shows the request landed, resolve the receipt as committed via POST /continuity/effects/{id}/resolve and attach the provider's receipt ID. If the provider shows it never applied, resolve it as abandoned and record what you looked at. If the evidence is inconclusive, leave it unknown and escalate.

External workers follow the same shape with one difference. When a step is handed to a worker over the REST long-poll API, the receipt stays Dispatched while the task is queued. The worker's completion callback commits the receipt before the task is marked complete, and the failure callback marks it Unknown before task retry handling runs. Callbacks from a worker that does not own the task are rejected, so a second worker claiming a leaked task cannot settle the first worker's receipt.

GuaranteeAchievable?Mechanism in Orch8
Exactly-once internal stateYesOne database transaction per transition
At-most-once external effect, per attemptYesReceipt written before dispatch, CAS on every transition
At-most-once external effect, per request contentYes, when declaredInvariant plus advisory lock and FOR UPDATE at dispatch
At-least-once external effectYesClean errors advance the attempt and get a fresh receipt
Exactly-once external effectNoNeeds a provider that honours idempotency keys

The bottom row is the one vendors blur. You can get close with a provider that honours idempotency keys, but that is the provider's guarantee. The engine carries the key in the receipt so a reconciliation can use it.

So the practical advice is to pass an idempotency key to every provider that accepts one, and let the engine keep it on the receipt. The engine's job is to make sure you never send twice without knowing, and to remember enough to find out what happened. The provider's job is to make a duplicate harmless if you do.

None of this requires Orch8. The parts that transfer to any service that calls someone else's API from inside retryable work:

  1. Write intent durably before you dispatch. If the record is not committed, you cannot later prove you tried.
  2. Derive the record's identity from its position in the work, not from a fresh UUID at call time. Retries must land on the same row.
  3. Give ambiguity its own state. "Sent, outcome unknown" is neither a failure nor a success, and modelling it as either causes a different bug.
  4. Make the state transition conditional at the storage layer. UPDATE ... WHERE state = expected is mutual exclusion. Check-then-act in application code is not.
  5. If two callers can race on the same content, serialise them in the database (an advisory lock, or a unique index on the content hash) rather than in memory.
  6. Classify handlers pessimistically. Anything you did not write is side-effecting until proven otherwise.
  7. Build the reconciliation path before you need it, and alert on the depth of the unresolved queue.

And say what you guarantee. "At-most-once, with an explicit unknown state and a reconciliation path" is a claim a reader can check. "Exactly-once" usually means something narrower that they will discover during an incident.

StateMeaningRetry allowed?How it resolves
PlannedDecided, nothing written yetYesAdvances to Prepared
PreparedReceipt durable, not yet sentYesAdvances to Dispatched
DispatchedSent, outcome pendingNoCommitted or Unknown
CommittedProvider confirmedNoTerminal, unless compensated
UnknownSent, never heard backNo, blocks the instanceVerified, Committed, Compensated, or Abandoned
VerifiedReconciled against providerNoTerminal, unless compensated
CompensatedUndone by compensating actionYes, new attemptTerminal
AbandonedOperator wrote it offNoTerminal, audited

Three of the eight states are operator or reconciler outcomes. Ambiguity is common enough to deserve first-class states rather than a paragraph in a runbook.

The pillar guide covers the full engine: snapshot execution, the crate architecture, storage backends, SDK design, and the trade-offs behind each decision.

Read the durable workflow engine architecture guide →