Snapshot vs Event Replay: Two Ways to Build a Durable Workflow Engine

Temporal rebuilds a workflow by re-running its code against recorded history. My engine, Orch8, reads the instance row and the list of finished steps and continues. That one decision sets your determinism rules, your versioning story, your payload limits, and what you can do when production breaks. Each model is better at something the other cannot do.

By Oleksii Vasylenko, Technical Lead · Published · Updated · 14 min read

Where this comes from. Orch8 is a durable workflow engine I built solo in Rust: every crate, every SDK, every integration. Ten-crate workspace, PostgreSQL and SQLite backends, server and on-device execution. The details below are the actual implementation, not a reference architecture.

A durable workflow engine has one job that ordinary code does not: when the process dies halfway through a workflow, the workflow has to continue from where it was, not from the beginning. There are two ways to build that. I spent a long time reading about one of them and then built the other, and this is the comparison I wish someone had written down before I chose.

Event replay. Persist an append-only log of everything that happened: this activity was scheduled, this timer fired, this result came back. On resume, run the workflow function again from the top. Every time it asks for something the log already has, hand it the recorded answer instead of doing the real thing. The function reaches the point where it left off and carries on. Temporal, Cadence, and Azure Durable Functions all work this way.

State snapshots. Persist the current state after each step completes. On resume, load that state and run the next step. Nothing is re-executed. This is what Orch8 does.

Event replay versus state snapshot resumeReplay loads the full history and re-executes the workflow function from the top. A snapshot engine reads one row and continues at the next step.State snapshots — resume is O(1)crashSELECT ... WHERE id = ?continue at next stepEvent replay — resume is O(n)crashload full historyre-execute workflow fnfrom the topfeed recorded resultsinstead of real callsreach the cut pointcontinue
The same crash handled two ways. Replay re-runs the workflow function against its history. A snapshot engine reads the saved state and continues at the next step.

Both work. Neither wins across the board, and the marketing on each side tends to describe the other side's worst case as its normal case. So before comparing them I want to be concrete about what a snapshot physically is in my engine, because "load the state" hides a couple of details that matter.

Orch8 vocabulary, briefly. A workflow definition is a sequence: a versioned JSON document made of blocks. The simplest block is a step, which names a handler (built-in ones like http_request or llm_call, or a handler you register) and its parameters. A run of a sequence is an instance. The engine is one Rust binary that runs a tick loop against Postgres, or SQLite for tests and on-device use.

An instance is one row in task_instances. The columns that matter for resume are state, next_fire_at, and context, a JSONB document holding the instance's input data and whatever steps have written into it. The outputs of individual steps live in a second table, block_outputs, one row per execution of a block. There is no event history table. There is nothing the engine has to walk in order.

CREATE TABLE task_instances (
    id              UUID PRIMARY KEY,
    sequence_id     UUID NOT NULL REFERENCES sequences(id),  -- pinned version
    tenant_id       TEXT NOT NULL,
    state           TEXT NOT NULL DEFAULT 'scheduled',
    next_fire_at    TIMESTAMPTZ,
    metadata        JSONB NOT NULL DEFAULT '{}',
    context         JSONB NOT NULL DEFAULT '{}',   -- the snapshot
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

CREATE TABLE block_outputs (
    instance_id     UUID NOT NULL REFERENCES task_instances(id),
    block_id        TEXT NOT NULL,
    output          JSONB,
    output_ref      TEXT,          -- externalised payload, or a sentinel
    attempt         SMALLINT NOT NULL DEFAULT 0,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
migrations/002_create_task_instances.sql and 004_create_block_outputs.sql, trimmed. Two tables are the whole snapshot.

When a step finishes, run_step_core in orch8-engine/src/handlers/step.rs inserts a block_outputs row for that block and attempt, then the scheduler bumps a step counter on the instance. If the step was the last one, the instance flips to completed. If the next block has a delay or is waiting on a signal, the row goes back to scheduled with a next_fire_at. That write is the checkpoint. A crash one millisecond later loses nothing that was committed.

Resume is the same path as first execution. Every tick, the scheduler claims due rows with FOR UPDATE SKIP LOCKED, then runs two batched queries for the whole claimed set: pending signals, and the list of block ids that already have a completed output. Then the step loop walks the sequence definition and skips anything in that list.

let mut completed_blocks = completed_block_ids;   // prefetched per batch
completed_blocks.sort_unstable();
completed_blocks.dedup();

for block in blocks {
    let BlockDefinition::Step(step_def) = block else { unreachable!() };

    if completed_blocks.binary_search(&step_def.id).is_ok() {
        continue;                                     // already ran, skip
    }

    let outcome = step_exec::execute_step_block(
        ctx.storage, ctx.handlers, ctx.webhook_config,
        ctx.externalize_threshold, instance, step_def,
        ctx.cancel, ctx.clock,
    ).await?;
    // Completed -> insert into completed_blocks, bump counter, carry on
    // Deferred / Failed -> return; the row was already updated
}
orch8-engine/src/scheduler.rs, execute_step_loop. The "resume" is a binary search over completed block ids; blocks already done are skipped.

So "loads a row" was a simplification I used to make in conversation, and it is not quite right. Resume reads the instance row, one batched query for completed block ids, and then individual block_outputs rows on demand when a later step references {{ outputs.some_step }} in its parameters. The cost scales with the number of blocks in the definition, not with how many things have happened to the instance. A loop that has iterated four thousand times has four thousand block_outputs rows, and the resume still asks for the latest one.

Replay only produces the right state if re-running the workflow function makes the same calls in the same order as the first run did. Temporal's documentation states it directly: the code must "make the same Workflow API calls in the same sequence, given the same input". During replay the commands the function emits are compared with the recorded history, and a mismatch raises a non-deterministic error. That means ordinary code is off the table inside a workflow function:

// Time: a replay hours later takes a different branch
if (Date.now() - startTime > TIMEOUT) { ... }

// Randomness: replay picks a different variant
const variant = Math.random() > 0.5 ? "a" : "b";

// Ambient state: the config changed since the original run
const limit = await redis.get("rate_limit");

// Iteration order: not guaranteed stable across runtimes
for (const key of Object.keys(payload)) { ... }
All four of these are bugs under replay, and the first three are silent until a replay happens to take the other branch.

Replay engines handle this properly. They supply deterministic replacements for time and randomness, they push anything with a side effect out into activities, and they detect non-determinism at replay time and fail loudly instead of corrupting state. It is a real solution. But it is a rule every developer on the team has to carry in their head, and a violation shows up at runtime, on resume, usually weeks after the code was written.

Snapshots have no determinism requirement because nothing is re-executed. Each step runs once, its output is persisted, and the next step reads that output. Date.now() is fine. Math.random() is fine. Reading config in the middle of a workflow is fine. There is no cleverness in this. The code that would need to be deterministic never runs a second time.

Why replay requires deterministic workflow code and snapshots do notRe-executing code on resume forces determinism constraints but enables replaying a production failure against new code. Snapshots allow ordinary code but give up the replay debugger.yes, replayno, snapshotworkflow codere-executedon resume?must be deterministicordinary code is fineno Date.now()no Math.random()no ambient readsstable iteration orderbuys: replay a productionfailure against new codecosts: no replay debugger,audit log only
Re-executing on resume imposes the four constraints on the left and enables the replay debugger on the right. A snapshot engine has neither.

And that debugger is the thing I miss. Because a replay engine can re-execute the workflow function against recorded history, you can take a production failure, pull its history, and replay it on your laptop against new code to see what the fix does. Orch8 cannot do that. There is no history to feed a re-execution. What I have instead is the saved state at each step, an append-only audit_log of state transitions, and step_logs per block. That is enough to diagnose most incidents and it is strictly less than replaying one.

Replay reconstructs state by re-running against the full history, so a workflow with 40,000 recorded events processes 40,000 events every time it wakes up. For a five-step workflow this does not matter. For a 90-day subscription lifecycle, a document approval that sits for weeks, or an agent loop making thousands of tool calls, it is a cost that grows on every wake-up. Temporal Cloud caps a single execution at 51,200 events or 50 MB of history, with a warning at 10,240 events.

The standard answer is continue-as-new: end the workflow periodically and start a fresh one that carries forward the state you still need. It works. It is also orchestration code your team writes and maintains for no reason other than the persistence model. In one private client codebase I reviewed, the continue-as-new handlers, state externalisation helpers, payload codecs, and search-attribute registration added up to roughly 2,400 lines. I cannot share that tree, so treat it as one data point.

Under snapshots, resume cost does not depend on how long the instance has been alive. The scheduler reads the instance row and a list of block ids, then skips the ones that are done. An instance on its four-thousandth loop iteration resumes at the same speed as one on its second. There is no continue-as-new because there is nothing to truncate.

The cost shows up somewhere else, and I want to name it. The snapshot is the only authoritative state. If a bug writes a bad context, there is no history to rebuild it from. Replay engines have a real durability advantage here: the log is the truth, the state is derived, and a state bug is fixed by correcting the derivation and replaying. What I have is the audit_log table, which records every state transition with the block id and details, and a checkpoints table that stores a JSONB copy of execution state on demand and is what fork-from-checkpoint uses. Those let me restore an instance to an earlier point. They do not let me recompute what it should have been.

Long-running workflows outlive the code that started them. You deploy on Tuesday, and an instance that started Monday is still running on Friday. Both models have to answer what happens to it, and neither answer is comfortable.

Under replay, changed workflow code means in-flight workflows replay their old history against the new function. If the new function makes a different call at some point, that is a non-determinism error. The fix is a version marker in the code (Temporal calls this patching, with GetVersion or Patched calls) that keeps the old path alive for old executions. Correct, and it accumulates. I have seen workflow functions with five layers of version branches that nobody dares delete because nobody is sure the last affected execution has finished.

Under snapshots, the instance row carries a sequence_id that points at one exact version of the definition. In-flight instances keep running on that version. New instances get the new one. No markers in code. What changes instead is that the state schema becomes the compatibility surface: if a new version changes the shape of what a step writes into context, an old instance that gets migrated forward carries data the new steps do not understand. Same problem, moved from code branches into data.

What happens to an in-flight workflow when you deploy new codeUnder replay, history replays against the new function and needs version markers. Under snapshots, instances stay pinned to a version and the risk moves to state schema compatibility.deploy new workflow versionin-flight instanceEvent replayState snapshotshistory replays againstthe NEW functiondivergence = non-determinismerrorfix: version markers in code(they accumulate forever)instance stays pinned toits sequence versionnew instances get the new onerisk moves to state schemacompatibility
Replay pushes the versioning problem into workflow code as version branches. Snapshots push it into the data as schema compatibility. Neither model removes it.

I put the safety net in the release path instead of in the workflow code. A release in Orch8 is a row in workflow_releases that pins a baseline sequence id and version and a candidate sequence id and version, and moves through draft, validating, ready, canary, and then promoted, paused, or rolled back. In-flight instances default to pin, meaning they stay on their version. Before a candidate can reach canary, three checks run on it.

The first is a semantic diff. It extracts a StepFacts record for every step in both versions and compares them field by field, and each difference gets a severity: informational, behavioral, side-effect risk, or incompatible.

struct StepFacts {
    handler: String,                  // changed handler on a side-effecting
    params: Value,                    //   step => SideEffectRisk
    queue: Option<String>,
    retry: Option<Value>,             // Behavioral
    timeout: Option<Value>,
    deadline: Option<Value>,
    delay: Option<Value>,
    send_window: Option<Value>,
    rate_limit_key: Option<String>,
    wait_for_input: bool,             // approval gate added or removed
    fallback_handler: Option<String>,
    output_schema: Option<Value>,     // downstream consumers may break
    when: Option<String>,             // guard change alters reachability
    compensation: Option<Value>,      // SideEffectRisk
}

pub enum DiffSeverity {
    Informational,   // cosmetic or metadata only
    Behavioral,      // paths, timing, retries change
    SideEffectRisk,  // an external effect may fire again or differently
    Incompatible,    // dangling references, narrowed inputs
}
orch8-engine/src/release_diff.rs. A changed handler on a side-effecting step is a different class of risk from a changed timeout.

The second is a dataflow compiler in orch8-engine/src/dataflow.rs. It collects every {{ outputs.x.y }} and {{ data.z }} reference in the candidate and checks it against the declared output_schema of the producing step and the sequence's input_schema. A reference to a step that no longer exists is MISSING_PRODUCER; a path the producer's schema does not contain is SCHEMA_PATH_MISSING. Both are treated as incompatible and block the release. The third is historical replay of real past instances against the candidate with side effects disabled, which is the closest I get to the replay debugger, and only works because step outputs are stored.

That is a lot of machinery. Replay engines get part of it for free, because non-determinism detection catches a subset of these problems at runtime without anybody writing a compiler. I built the compiler because I did not have the runtime check. One CLI command, orch8 release gate <id>, folds the three checks into a single exit code for CI.

Replay engines cap the size of history because the whole history travels between server and worker on every replay. On Temporal Cloud a single request payload is limited to 2 MB and history to 50 MB. The workaround is a payload codec that swaps large values for references to external storage. That is a good pattern and it is also one more thing to run.

In Orch8 the state is a JSONB column, so the hard limit is whatever Postgres allows. I still bound it. max_context_bytes defaults to 256 KiB, and fields above externalize_output_threshold (off by default, configurable) are moved to an externalized_state table with optional zstd compression and an expires_at, leaving a marker in the context. When the scheduler claims a batch it hydrates those markers in one query. The reason is query performance, not protocol. A 40 MB JSONB column ruins every query that touches the table.

The quieter advantage of snapshots is that state is queryable with SQL. Under replay, the state of a running workflow lives inside the workflow function's memory while it executes, so answering "which workflows are stuck on legal review" means writing a query handler per workflow and deploying it. Under snapshots it is a query against a GIN-indexed metadata column.

-- Every instance waiting on a specific approval, across all sequences
SELECT id, sequence_id, context->'data'->>'order_id'
FROM task_instances
WHERE state = 'waiting'
  AND metadata @> '{"awaiting": "legal_review"}';

-- Backlog by tenant, right now
SELECT tenant_id, count(*)
FROM task_instances
WHERE state = 'scheduled' AND next_fire_at <= now()
GROUP BY tenant_id ORDER BY 2 DESC;

-- Latest output of one step for one instance
SELECT output FROM block_outputs
WHERE instance_id = $1 AND block_id = 'charge_card'
ORDER BY created_at DESC, id DESC LIMIT 1;
Operational questions I have run during incidents without deploying anything. The GIN index on metadata is in migrations/009_create_indexes.sql.

The questions you need at 3am are rarely the ones you wrote a query handler for. I have leaned on this more than any feature on a comparison chart.

Temporal's production topology is a frontend service, a history service, a matching service, and worker services, plus Cassandra or Postgres, plus Elasticsearch for visibility, plus your own workers. Every piece is there for a reason. That architecture is what lets it run enormous workloads with strong isolation between namespaces.

It is also a full-time job. On a team of five, adopting it means one person becomes the workflow-infrastructure person. That cost is real and it does not appear in a feature matrix.

Orch8 is one binary and Postgres. The same binary runs in five roles (all-in-one, control, executor, gateway, edge), so scaling out is a topology change rather than a new set of services. SQLite runs the identical engine for tests, embedded use, and on-device execution, which means the test suite needs no running server.

The counterpoint: Temporal has years of production hardening across thousands of deployments, a large ecosystem, and people you can hire who already know it. "One binary" is an operational advantage and an ecosystem disadvantage. Which one dominates depends on team size and how much risk you are willing to carry yourself. Below roughly ten engineers I think operational simplicity wins. Above that, the ecosystem often does.

Neither model is correct in general. Each is correct for particular shapes of problem, and the shape is usually knowable before you start.

  • Event replay if you need to replay production failures against new code on your laptop, and your workflows are short enough that history stays small. It also fits teams that can run a multi-service cluster, or that value the hiring pool.
  • Snapshots if workflows run for weeks or months, developers need to write ordinary non-deterministic code, you want to query state with SQL, or the engine has to run on a device. It also fits a team whose operational budget is one binary plus Postgres.
  • Neither if a cron job and an idempotent script would do. A durable engine earns its place only when steps have to survive failures, wait on humans, or coordinate across systems. A retry loop and a status column solve a surprising share of what people reach for a workflow engine to solve.

Whichever you pick, distrust any description that shows one model with no downsides. Replay buys a real debugger and pays in determinism rules and history cost. Snapshots buy freedom in workflow code and constant-time resume, and pay in weaker reconstruction and release safety you have to build yourself. Those are the actual trades, and I have felt both sides of them.

DimensionEvent replayState snapshots
Resume costProportional to history lengthInstance row plus completed block ids
DeterminismRequired in workflow codeNot required
Long-running workflowsNeeds continue-as-new plumbingNo special handling
Failure reconstructionReplay history against new code locallyAudit log, step logs, checkpoints
State visibilityQuery handlers per workflowSQL over JSONB
VersioningVersion markers accumulate in codePinned versions plus schema compatibility
Payload limitsPer-payload and history caps, codec for large dataConfigurable, externalised above a threshold
Corrupted state recoveryRebuild from the logRestore from a checkpoint
Operational footprintMulti-service clusterOne binary plus Postgres or SQLite
Ecosystem and hiringMature, largeSmall

Rows four, eight, and ten favour replay. The rest favour snapshots. Which rows matter depends on the workflows you actually run.

Exactly-Once Side Effects →

What either execution model can promise about effects on external systems, and the receipt that decides whether a crashed step may re-run.

The pillar guide covers the crate architecture, storage backends, SDK design, and the reasoning behind the snapshot execution model.

Read the durable workflow engine architecture guide →