Snapshot vs Event Replay: Two Ways to Build a Durable Workflow Engine
Temporal rebuilds a workflow by re-running its code against recorded history. My engine, Orch8, reads the instance row and the list of finished steps and continues. That one decision sets your determinism rules, your versioning story, your payload limits, and what you can do when production breaks. Each model is better at something the other cannot do.
By Oleksii Vasylenko, Technical Lead · Published · Updated · 14 min read
Where this comes from. Orch8 is a durable workflow engine I built solo in Rust: every crate, every SDK, every integration. Ten-crate workspace, PostgreSQL and SQLite backends, server and on-device execution. The details below are the actual implementation, not a reference architecture.
Two ways to survive a crash
A durable workflow engine has one job that ordinary code does not: when the process dies halfway through a workflow, the workflow has to continue from where it was, not from the beginning. There are two ways to build that. I spent a long time reading about one of them and then built the other, and this is the comparison I wish someone had written down before I chose.
Event replay. Persist an append-only log of everything that happened: this activity was scheduled, this timer fired, this result came back. On resume, run the workflow function again from the top. Every time it asks for something the log already has, hand it the recorded answer instead of doing the real thing. The function reaches the point where it left off and carries on. Temporal, Cadence, and Azure Durable Functions all work this way.
State snapshots. Persist the current state after each step completes. On resume, load that state and run the next step. Nothing is re-executed. This is what Orch8 does.
Both work. Neither wins across the board, and the marketing on each side tends to describe the other side's worst case as its normal case. So before comparing them I want to be concrete about what a snapshot physically is in my engine, because "load the state" hides a couple of details that matter.
What a snapshot is, concretely
Orch8 vocabulary, briefly. A workflow definition is a sequence: a versioned JSON document made of blocks. The simplest block is a step, which names a handler (built-in ones like http_request or llm_call, or a handler you register) and its parameters. A run of a sequence is an instance. The engine is one Rust binary that runs a tick loop against Postgres, or SQLite for tests and on-device use.
An instance is one row in task_instances. The columns that matter for resume are state, next_fire_at, and context, a JSONB document holding the instance's input data and whatever steps have written into it. The outputs of individual steps live in a second table, block_outputs, one row per execution of a block. There is no event history table. There is nothing the engine has to walk in order.
CREATE TABLE task_instances (
id UUID PRIMARY KEY,
sequence_id UUID NOT NULL REFERENCES sequences(id), -- pinned version
tenant_id TEXT NOT NULL,
state TEXT NOT NULL DEFAULT 'scheduled',
next_fire_at TIMESTAMPTZ,
metadata JSONB NOT NULL DEFAULT '{}',
context JSONB NOT NULL DEFAULT '{}', -- the snapshot
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE TABLE block_outputs (
instance_id UUID NOT NULL REFERENCES task_instances(id),
block_id TEXT NOT NULL,
output JSONB,
output_ref TEXT, -- externalised payload, or a sentinel
attempt SMALLINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);When a step finishes, run_step_core in orch8-engine/src/handlers/step.rs inserts a block_outputs row for that block and attempt, then the scheduler bumps a step counter on the instance. If the step was the last one, the instance flips to completed. If the next block has a delay or is waiting on a signal, the row goes back to scheduled with a next_fire_at. That write is the checkpoint. A crash one millisecond later loses nothing that was committed.
Resume is the same path as first execution. Every tick, the scheduler claims due rows with FOR UPDATE SKIP LOCKED, then runs two batched queries for the whole claimed set: pending signals, and the list of block ids that already have a completed output. Then the step loop walks the sequence definition and skips anything in that list.
let mut completed_blocks = completed_block_ids; // prefetched per batch
completed_blocks.sort_unstable();
completed_blocks.dedup();
for block in blocks {
let BlockDefinition::Step(step_def) = block else { unreachable!() };
if completed_blocks.binary_search(&step_def.id).is_ok() {
continue; // already ran, skip
}
let outcome = step_exec::execute_step_block(
ctx.storage, ctx.handlers, ctx.webhook_config,
ctx.externalize_threshold, instance, step_def,
ctx.cancel, ctx.clock,
).await?;
// Completed -> insert into completed_blocks, bump counter, carry on
// Deferred / Failed -> return; the row was already updated
}So "loads a row" was a simplification I used to make in conversation, and it is not quite right. Resume reads the instance row, one batched query for completed block ids, and then individual block_outputs rows on demand when a later step references {{ outputs.some_step }} in its parameters. The cost scales with the number of blocks in the definition, not with how many things have happened to the instance. A loop that has iterated four thousand times has four thousand block_outputs rows, and the resume still asks for the latest one.
The determinism tax of event replay
Replay only produces the right state if re-running the workflow function makes the same calls in the same order as the first run did. Temporal's documentation states it directly: the code must "make the same Workflow API calls in the same sequence, given the same input". During replay the commands the function emits are compared with the recorded history, and a mismatch raises a non-deterministic error. That means ordinary code is off the table inside a workflow function:
// Time: a replay hours later takes a different branch
if (Date.now() - startTime > TIMEOUT) { ... }
// Randomness: replay picks a different variant
const variant = Math.random() > 0.5 ? "a" : "b";
// Ambient state: the config changed since the original run
const limit = await redis.get("rate_limit");
// Iteration order: not guaranteed stable across runtimes
for (const key of Object.keys(payload)) { ... }Replay engines handle this properly. They supply deterministic replacements for time and randomness, they push anything with a side effect out into activities, and they detect non-determinism at replay time and fail loudly instead of corrupting state. It is a real solution. But it is a rule every developer on the team has to carry in their head, and a violation shows up at runtime, on resume, usually weeks after the code was written.
Snapshots have no determinism requirement because nothing is re-executed. Each step runs once, its output is persisted, and the next step reads that output. Date.now() is fine. Math.random() is fine. Reading config in the middle of a workflow is fine. There is no cleverness in this. The code that would need to be deterministic never runs a second time.
And that debugger is the thing I miss. Because a replay engine can re-execute the workflow function against recorded history, you can take a production failure, pull its history, and replay it on your laptop against new code to see what the fix does. Orch8 cannot do that. There is no history to feed a re-execution. What I have instead is the saved state at each step, an append-only audit_log of state transitions, and step_logs per block. That is enough to diagnose most incidents and it is strictly less than replaying one.
Resume cost and history growth
Replay reconstructs state by re-running against the full history, so a workflow with 40,000 recorded events processes 40,000 events every time it wakes up. For a five-step workflow this does not matter. For a 90-day subscription lifecycle, a document approval that sits for weeks, or an agent loop making thousands of tool calls, it is a cost that grows on every wake-up. Temporal Cloud caps a single execution at 51,200 events or 50 MB of history, with a warning at 10,240 events.
The standard answer is continue-as-new: end the workflow periodically and start a fresh one that carries forward the state you still need. It works. It is also orchestration code your team writes and maintains for no reason other than the persistence model. In one private client codebase I reviewed, the continue-as-new handlers, state externalisation helpers, payload codecs, and search-attribute registration added up to roughly 2,400 lines. I cannot share that tree, so treat it as one data point.
Under snapshots, resume cost does not depend on how long the instance has been alive. The scheduler reads the instance row and a list of block ids, then skips the ones that are done. An instance on its four-thousandth loop iteration resumes at the same speed as one on its second. There is no continue-as-new because there is nothing to truncate.
The cost shows up somewhere else, and I want to name it. The snapshot is the only authoritative state. If a bug writes a bad context, there is no history to rebuild it from. Replay engines have a real durability advantage here: the log is the truth, the state is derived, and a state bug is fixed by correcting the derivation and replaying. What I have is the audit_log table, which records every state transition with the block id and details, and a checkpoints table that stores a JSONB copy of execution state on demand and is what fork-from-checkpoint uses. Those let me restore an instance to an earlier point. They do not let me recompute what it should have been.
Versioning: deploying new code while workflows are running
Long-running workflows outlive the code that started them. You deploy on Tuesday, and an instance that started Monday is still running on Friday. Both models have to answer what happens to it, and neither answer is comfortable.
Under replay, changed workflow code means in-flight workflows replay their old history against the new function. If the new function makes a different call at some point, that is a non-determinism error. The fix is a version marker in the code (Temporal calls this patching, with GetVersion or Patched calls) that keeps the old path alive for old executions. Correct, and it accumulates. I have seen workflow functions with five layers of version branches that nobody dares delete because nobody is sure the last affected execution has finished.
Under snapshots, the instance row carries a sequence_id that points at one exact version of the definition. In-flight instances keep running on that version. New instances get the new one. No markers in code. What changes instead is that the state schema becomes the compatibility surface: if a new version changes the shape of what a step writes into context, an old instance that gets migrated forward carries data the new steps do not understand. Same problem, moved from code branches into data.
I put the safety net in the release path instead of in the workflow code. A release in Orch8 is a row in workflow_releases that pins a baseline sequence id and version and a candidate sequence id and version, and moves through draft, validating, ready, canary, and then promoted, paused, or rolled back. In-flight instances default to pin, meaning they stay on their version. Before a candidate can reach canary, three checks run on it.
The first is a semantic diff. It extracts a StepFacts record for every step in both versions and compares them field by field, and each difference gets a severity: informational, behavioral, side-effect risk, or incompatible.
struct StepFacts {
handler: String, // changed handler on a side-effecting
params: Value, // step => SideEffectRisk
queue: Option<String>,
retry: Option<Value>, // Behavioral
timeout: Option<Value>,
deadline: Option<Value>,
delay: Option<Value>,
send_window: Option<Value>,
rate_limit_key: Option<String>,
wait_for_input: bool, // approval gate added or removed
fallback_handler: Option<String>,
output_schema: Option<Value>, // downstream consumers may break
when: Option<String>, // guard change alters reachability
compensation: Option<Value>, // SideEffectRisk
}
pub enum DiffSeverity {
Informational, // cosmetic or metadata only
Behavioral, // paths, timing, retries change
SideEffectRisk, // an external effect may fire again or differently
Incompatible, // dangling references, narrowed inputs
}The second is a dataflow compiler in orch8-engine/src/dataflow.rs. It collects every {{ outputs.x.y }} and {{ data.z }} reference in the candidate and checks it against the declared output_schema of the producing step and the sequence's input_schema. A reference to a step that no longer exists is MISSING_PRODUCER; a path the producer's schema does not contain is SCHEMA_PATH_MISSING. Both are treated as incompatible and block the release. The third is historical replay of real past instances against the candidate with side effects disabled, which is the closest I get to the replay debugger, and only works because step outputs are stored.
That is a lot of machinery. Replay engines get part of it for free, because non-determinism detection catches a subset of these problems at runtime without anybody writing a compiler. I built the compiler because I did not have the runtime check. One CLI command, orch8 release gate <id>, folds the three checks into a single exit code for CI.
Payload limits and what the database can see
Replay engines cap the size of history because the whole history travels between server and worker on every replay. On Temporal Cloud a single request payload is limited to 2 MB and history to 50 MB. The workaround is a payload codec that swaps large values for references to external storage. That is a good pattern and it is also one more thing to run.
In Orch8 the state is a JSONB column, so the hard limit is whatever Postgres allows. I still bound it. max_context_bytes defaults to 256 KiB, and fields above externalize_output_threshold (off by default, configurable) are moved to an externalized_state table with optional zstd compression and an expires_at, leaving a marker in the context. When the scheduler claims a batch it hydrates those markers in one query. The reason is query performance, not protocol. A 40 MB JSONB column ruins every query that touches the table.
The quieter advantage of snapshots is that state is queryable with SQL. Under replay, the state of a running workflow lives inside the workflow function's memory while it executes, so answering "which workflows are stuck on legal review" means writing a query handler per workflow and deploying it. Under snapshots it is a query against a GIN-indexed metadata column.
-- Every instance waiting on a specific approval, across all sequences
SELECT id, sequence_id, context->'data'->>'order_id'
FROM task_instances
WHERE state = 'waiting'
AND metadata @> '{"awaiting": "legal_review"}';
-- Backlog by tenant, right now
SELECT tenant_id, count(*)
FROM task_instances
WHERE state = 'scheduled' AND next_fire_at <= now()
GROUP BY tenant_id ORDER BY 2 DESC;
-- Latest output of one step for one instance
SELECT output FROM block_outputs
WHERE instance_id = $1 AND block_id = 'charge_card'
ORDER BY created_at DESC, id DESC LIMIT 1;The questions you need at 3am are rarely the ones you wrote a query handler for. I have leaned on this more than any feature on a comparison chart.
Operational footprint
Temporal's production topology is a frontend service, a history service, a matching service, and worker services, plus Cassandra or Postgres, plus Elasticsearch for visibility, plus your own workers. Every piece is there for a reason. That architecture is what lets it run enormous workloads with strong isolation between namespaces.
It is also a full-time job. On a team of five, adopting it means one person becomes the workflow-infrastructure person. That cost is real and it does not appear in a feature matrix.
Orch8 is one binary and Postgres. The same binary runs in five roles (all-in-one, control, executor, gateway, edge), so scaling out is a topology change rather than a new set of services. SQLite runs the identical engine for tests, embedded use, and on-device execution, which means the test suite needs no running server.
The counterpoint: Temporal has years of production hardening across thousands of deployments, a large ecosystem, and people you can hire who already know it. "One binary" is an operational advantage and an ecosystem disadvantage. Which one dominates depends on team size and how much risk you are willing to carry yourself. Below roughly ten engineers I think operational simplicity wins. Above that, the ecosystem often does.
How to choose between replay and snapshots
Neither model is correct in general. Each is correct for particular shapes of problem, and the shape is usually knowable before you start.
- Event replay if you need to replay production failures against new code on your laptop, and your workflows are short enough that history stays small. It also fits teams that can run a multi-service cluster, or that value the hiring pool.
- Snapshots if workflows run for weeks or months, developers need to write ordinary non-deterministic code, you want to query state with SQL, or the engine has to run on a device. It also fits a team whose operational budget is one binary plus Postgres.
- Neither if a cron job and an idempotent script would do. A durable engine earns its place only when steps have to survive failures, wait on humans, or coordinate across systems. A retry loop and a status column solve a surprising share of what people reach for a workflow engine to solve.
Whichever you pick, distrust any description that shows one model with no downsides. Replay buys a real debugger and pays in determinism rules and history cost. Snapshots buy freedom in workflow code and constant-time resume, and pay in weaker reconstruction and release safety you have to build yourself. Those are the actual trades, and I have felt both sides of them.
Event Replay vs State Snapshots
| Dimension | Event replay | State snapshots |
|---|---|---|
| Resume cost | Proportional to history length | Instance row plus completed block ids |
| Determinism | Required in workflow code | Not required |
| Long-running workflows | Needs continue-as-new plumbing | No special handling |
| Failure reconstruction | Replay history against new code locally | Audit log, step logs, checkpoints |
| State visibility | Query handlers per workflow | SQL over JSONB |
| Versioning | Version markers accumulate in code | Pinned versions plus schema compatibility |
| Payload limits | Per-payload and history caps, codec for large data | Configurable, externalised above a threshold |
| Corrupted state recovery | Rebuild from the log | Restore from a checkpoint |
| Operational footprint | Multi-service cluster | One binary plus Postgres or SQLite |
| Ecosystem and hiring | Mature, large | Small |
Rows four, eight, and ten favour replay. The rest favour snapshots. Which rows matter depends on the workflows you actually run.
Related Matching Engine Guides
What either execution model can promise about effects on external systems, and the receipt that decides whether a crashed step may re-run.
The claim query and tick loop that resume snapshot state without a message broker.
What a portable state snapshot makes possible that replay history cannot.
Tenant isolation in a single-binary engine, from breakers to storage routing.
Related reading
The pillar guide covers the crate architecture, storage backends, SDK design, and the reasoning behind the snapshot execution model.
Read the durable workflow engine architecture guide →