Moving a Running Workflow Across a Trust Boundary
Copying a stopped process is a file copy. Moving a running one means transferring authority over its side effects. This is how Orch8 does that: an ownership row with an epoch, a signed and encrypted capsule, effect receipts keyed to the logical execution, and a provenance chain that anyone can verify.
By Oleksii Vasylenko, Technical Lead · Published · Updated · 16 min read
Where this comes from, and what state it is in. Orch8 is a durable workflow engine I built solo in Rust. The ownership model, capsule format, effect receipts, provenance chain, and federation envelope described here are real code with unit and integration tests, and the server-to-device roundtrip runs in the mobile test suite. What has not happened yet is a handoff across a live trust boundary between two organisations at production scale. Read this as a design I can defend line by line, not as a system with production incidents behind it.
Why move a running workflow at all
Portable workflow execution means taking a workflow that is partway through its steps and continuing it on a different machine, possibly one owned by someone else. Three situations pushed me into building it, and restarting the workflow from scratch on the new machine does not solve any of them.
- A phone goes offline in the middle of an onboarding flow the server was driving. The remaining steps should run on the device and reconcile when the network comes back.
- One step has to touch data that cannot leave a customer's own infrastructure or a jurisdiction. Only that step needs to run there. The rest of the orchestration does not.
- A process spans two companies. Neither will run the other's code or open its database, but the process is one process and needs one continuous record.
My first attempt was the obvious one: serialise the instance, send it, deserialise it, carry on. It works until the source does not actually stop. Then two runtimes both think they own the execution, and both call the payment API. Everything in this article exists to make that state unreachable, and most of it is checked by the database rather than by application code.
What Orch8 stores for a running workflow
Some vocabulary first, because the rest of the article uses it. In Orch8 a workflow definition is a sequence made of blocks. Each block names a handler, either a built-in such as http_request or one a user registered. A run of a sequence is an instance: one row in task_instances with a state, a JSONB context, and a next_fire_at timestamp. The engine is one Rust binary that runs a tick loop against Postgres or SQLite, claims due instances, and executes their next block. That is the whole engine for a workflow that never leaves its server.
Portable continuity adds a second identity on top of the instance. An instance ID is local to one runtime and changes every time the execution moves. The thing that stays constant is the continuity_id, and its authoritative record is one row in continuity_executions:
pub struct ContinuityExecution {
pub continuity_id: ContinuityId, // stable across every move
pub tenant_id: TenantId,
pub current_instance_id: InstanceId, // the local run on whichever runtime owns it now
pub owner_runtime_id: RuntimeId, // who may act right now
pub epoch: ExecutionEpoch, // u64, increments on every ownership change
pub state: OwnershipState, // Owned | Transferring | Completed
pub updated_at: DateTime<Utc>,
}
CREATE TABLE continuity_executions (
continuity_id UUID PRIMARY KEY,
tenant_id TEXT NOT NULL,
epoch BIGINT NOT NULL CHECK (epoch >= 0),
owner_runtime_id UUID NOT NULL,
state TEXT NOT NULL,
record JSONB NOT NULL,
updated_at TIMESTAMPTZ NOT NULL,
UNIQUE (tenant_id, continuity_id)
);Alongside it sit five more tables from the same migration: execution_handoffs (one row per transfer attempt, with its own version counter), execution_capsules (signed manifests with an expiry), runtime_capabilities (what each registered runtime says it can do, with an expiry), effect_receipts (every external side effect, keyed to the continuity ID rather than the instance), and provenance_entries (the hash chain). A later migration, 061, adds a unique index on (tenant_id, current_instance_id), so a local instance can belong to at most one continuity execution. I would rather have that as an index than as a check in Rust that a future code path can skip.
One handoff, step by step
The API surface is in orch8-api/src/continuity.rs and the transition rules in orch8-engine/src/continuity.rs. A server-to-server handoff inside one deployment goes through six calls. Every one of them is a compare-and-swap against a row it read a moment earlier.
- Preview.
POST /continuity/executions/{id}/handoff/previewnames a destination runtime and the capsule requirements.build_handoff_previewloads the destination's live capabilities, runsassess_compatibility, runs the placement policy, and lists every effect receipt that is not resolved. It returns apreview_sha256over all of that evidence. If any effect receipt is stillDispatchedorUnknown, the preview reportscompatible: false. You cannot move an execution while it has a side effect in flight whose outcome nobody knows. - Create.
POST .../handoffmust carry that digest.create_handoffrebuilds the preview and refuses if the digest no longer matches, so a handoff can only be created against evidence the caller actually looked at. The handoff row starts inRequestedwithversion = 0. - Export.
export_handoffchecks that the source instance isPausedorWaiting, moves the handoff toQuiescingwith a CAS on its version, builds the capsule (next section), and then callscommit_handoff_export. That single storage transaction flips the handoff toExportedand the execution fromOwnedtoTransferring, withWHERE epoch = $expected AND owner_runtime_id = $expected AND state = 'owned'on the update. If zero rows change, the whole thing rolls back and the caller getsStaleClaim. - Import. The destination calls
import_capsulewith the signed manifest and the encrypted payload.verify_and_import_paused_capsule_byteschecks the manifest signature against the engine's key, the payload length and SHA-256 against the manifest, decrypts, validates bounds, and creates a new local instance inPaused. Nothing about ownership changes yet. - Accept.
accept_handoffcomputesepoch.checked_next()and writes the new owner. The storage call updatescontinuity_executionsonly where the epoch, owner, andstate = 'transferring'still match, and only ifis_capsule_import_instanceconfirms the destination instance was created from this handoff's capsule. Then it appends aruntime_claimedprovenance entry. - Resume and complete.
resume_handoffmoves the destination instance toScheduledso the tick loop picks it up, and the handoff row walksAcceptedtoResumedtoCompleted. TheHandoffState::can_transition_totable has no path fromRequestedtoResumed, and no path out of a terminal state, so a client cannot skip the middle.
The mobile path is the same protocol with fewer round trips. orch8-mobile/src/continuity.rs exposes export_capsule, import_capsule, and activate_capsule. The device signs the manifest with its own Ed25519 key (the server learns that public key at registration through capsule_signing_public_key), and activate_capsule does the ownership CAS locally with cas_continuity_owner. The test portable_roundtrip_survives_redelivery_and_activation runs an export, imports the same capsule twice, activates it, and checks that the second import is a no-op rather than a second owner.
Ownership is an epoch, not a flag
A boolean is_owner cannot survive a network partition. The old owner still has its flag set and no way to learn that the world moved on. What works is a generation number that increases on every transfer, so that a write from a previous owner is recognisably from the past rather than merely late. This is the fencing-token idea from Martin Kleppmann's post on distributed locking, applied to execution authority instead of a lock.
pub struct ExecutionEpoch(u64);
impl ExecutionEpoch {
pub fn checked_next(self) -> Result<Self, ContinuityError> {
self.0
.checked_add(1)
.map(Self)
.ok_or(ContinuityError::EpochOverflow)
}
}The Transferring state is the second half of this. Between the export commit and the accept, neither runtime may dispatch effects: the source has given up Owned, and the destination has not yet claimed it. If the process dies in that window, the execution is parked in a state that says so. An operator can revoke the handoff (Exported to Revoked is a legal transition) and the source reclaims it at a fresh epoch. Without that intermediate state the failure mode is two rows of evidence that both say "mine".
What a capsule physically contains
The capsule is what actually crosses the wire. It has two parts: a signed manifest that travels as JSON, and an encrypted payload that is stored as an artifact and fetched or attached separately. export_paused_capsule_manifest in orch8-engine/src/capsule.rs assembles both.
pub struct CapsulePayload {
pub instance: CapsuleInstanceState, // sequence_id, namespace, priority, timezone,
// metadata, context, budget, parent_instance_id
pub checkpoint: Checkpoint, // the latest durable checkpoint
pub outputs: Vec<BlockOutput>, // every completed block's output
pub pending_waits: Vec<Value>,
pub pending_signals: Vec<Value>,
pub effect_ids: Vec<EffectId>, // receipts that belong to this execution
pub artifacts: Vec<ArtifactReference>,
pub stream_cursors: Vec<StreamCursor>,
pub redacted_audit_context: Value,
}
// MAX_OUTPUTS 10_000, MAX_PENDING_ITEMS 10_000, MAX_ENCODED_BYTES 64 MiB
pub struct CapsuleManifest {
pub schema: CapsuleSchemaVersion, // { major, minor }
pub capsule_id: CapsuleId,
pub continuity_id: ContinuityId,
pub source_instance_id: InstanceId,
pub epoch: ExecutionEpoch,
pub tenant_id: TenantId,
pub source_runtime_id: RuntimeId,
pub allowed_destination_runtime_id: Option<RuntimeId>,
pub sequence: SequenceIdentity, // id, version, content_sha256
pub checkpoint: CheckpointIdentity, // block_id, sha256
pub requirements_sha256: String,
pub payload_artifact: ArtifactReference, // key, sha256, bytes
pub provenance_head: Option<String>,
pub issued_at: DateTime<Utc>,
pub expires_at: DateTime<Utc>, // default 300 s, at most 3600 s
pub signing_key_id: String,
pub encryption_key_id: String,
}The payload is canonical JSON, encrypted with the engine's field encryptor using additional authenticated data made of the tenant, the continuity ID, the epoch, and the destination runtime. So a payload cannot be re-attached to a different manifest, replayed into a different tenant, or delivered to a runtime other than the one it was exported for, even by someone who has the ciphertext. For a device transfer the caller can supply a one-time transfer key instead of the engine key, which is what the mobile path does.
Export refuses unless the source instance is Paused or Waiting. A block that is mid-execution has no clean boundary to cut at, and the checkpoint the manifest names would not describe the state that actually left. The manifest also records the sequence content hash and the provenance head at export time, so the destination can tell whether it is looking at the same workflow definition and the same history the source had.
Effect receipts have to survive the move
This is where the naive copy fails hardest. Say the workflow charged a card on the server and then relocated to a phone. If the record of that charge does not travel, or travels but is not the record the phone consults, the phone re-runs the step. The at-most-once design in the effect receipts article only works if the receipt is keyed to something that survives relocation.
So an EffectReceipt carries continuity_id and epoch, and only secondarily instance_id. The receipts table has a foreign key to continuity_executions, not to the instance. When the execution lands on a new runtime with a new local instance, the receipts still belong to the same logical execution, and the capsule payload lists their IDs so the destination knows which ones to expect.
match storage.dispatch_effect_receipt_at_most_once(&tenant, &updated).await? {
EffectDispatchOutcome::Dispatched => { /* this caller owns the effect */ }
EffectDispatchOutcome::Duplicate => Err(EngineError::InvariantViolation { .. }),
EffectDispatchOutcome::Stale => Err(blocked(¤t)), // older epoch tried to act
}Stale and Duplicate are separate on purpose. A duplicate means the same execution tried to dispatch one effect twice, which is a workflow or engine bug. A stale outcome means a runtime from a superseded epoch tried to act after ownership moved, which is a handoff race. The two need different responses and different people looking at them.
And, as mentioned in the walk-through, the handoff preview refuses to report a destination as compatible while any receipt is unresolved. A Dispatched receipt with no answer yet, or an Unknown one waiting for reconciliation, blocks the move. I would rather delay a transfer than transfer an execution whose recent past nobody can vouch for.
The destination has to prove it can run it
A workflow that needs an S3 credential and a GPU cannot run on a phone. Finding that out after accept means an execution stranded on a runtime that can never advance it, and since ownership already moved, the origin cannot take it back without another handoff. So the requirements travel with the request and are checked before anything changes.
pub struct CapsuleRequirements {
pub handlers: Vec<String>, // every block handler must be registered there
pub plugins: Vec<String>,
pub credentials: Vec<String>,
pub regions: Vec<String>, // residency: destination must be in one of them
pub hardware: Vec<String>,
pub requires_network: bool,
pub requires_human_ui: bool, // only Mobile, Desktop, Browser can satisfy this
pub minimum_trust: Option<RuntimeTrustLevel>, // Unverified < Registered < Signed < Attested
}
pub enum RuntimeKind { Server, Edge, Mobile, Desktop, Browser }Runtimes register their capabilities with an expiry of at most five minutes, so a stale advertisement ages out on its own. Two functions read those facts. assess_compatibility produces findings with codes like HANDLERS_MISSING, REGION_NOT_ALLOWED, and TRUST_TOO_LOW, and it distinguishes Fail from Unknown: a runtime that did not report its connectivity gets NETWORK_UNKNOWN, which shows up in the preview. CapsuleRequirements::is_satisfied_by is the one used when a runtime actually claims work, and it treats unknown the same as failed. A preview can tell you "we are not sure". A claim cannot proceed on "not sure".
Delegating a slice of work to a device adds more checks on top: the current ownership epoch, live same-tenant registrations for both source and destination, every handler the isolated sub-sequence needs, a signed continuation grant bound to that destination, and a one-time token whose SHA-256 is stored and consumed on use. The device gets an isolated sub-sequence, not the whole execution, so a compromised phone can only reach its own slice.
Provenance: a hash chain with an external anchor
Once an execution has crossed three runtimes and two organisations, "what happened?" needs an answer that does not depend on trusting whoever is answering. Each boundary (capsule exported, runtime claimed, effect resolved, federation envelope signed) appends a ProvenanceEntry whose hash commits to the previous entry's hash.
pub struct ProvenanceEntry {
pub continuity_id: ContinuityId,
pub epoch: ExecutionEpoch,
pub kind: String, // "capsule_exported", "runtime_claimed", ...
pub redacted_summary: Option<String>, // operator-safe, bounded
pub payload_sha256: String,
pub previous_sha256: Option<String>, // the chain link
pub entry_sha256: String,
pub signing_key_id: Option<String>,
pub signature: Option<String>, // Ed25519 over entry_sha256
pub created_at: DateTime<Utc>,
}The entry hash is SHA-256 over a domain separator (orch8-provenance-v2), the continuity ID, the big-endian epoch, and length-framed copies of the kind, summary, payload digest, and previous hash. The length framing is there so that two different field splits cannot produce the same bytes. Storing digests instead of payloads is what makes the chain safe to retain and safe to show a counterparty: a regulator can verify that a decision was recorded and never edited without the audit log becoming a second copy of the customer data.
Key rotation uses a registry rather than re-signing history. The active continuity signing key is trusted automatically; retired key IDs map to their public keys through ORCH8_CONTINUITY_TRUSTED_SIGNING_KEYS_JSON, and verify_provenance_chain_with_keys looks each entry's signing_key_id up there. An entry signed by a key that is not in the registry fails with PROVENANCE_SIGNING_KEY_UNTRUSTED. Without that map, the first rotation would invalidate every entry signed before it.
Crossing into another organisation
Federation is where the trust assumptions change. Inside one deployment every runtime is mine. Across a federation boundary the destination belongs to someone else, and the design has to assume it might be wrong or hostile.
Peers come from operator configuration only, through ORCH8_FEDERATION_PEERS. They are never discovered and never self-register. parse_federation_peers in orch8-server/src/main.rs caps the list at 128 peers with unique IDs, requires an https:// endpoint under 2,048 bytes, a tenant allowlist of 1 to 256 entries, and a 64-character hex trust_root_sha256. That digest is the SHA-256 of the peer's Ed25519 public key bytes, and the verifier recomputes it before trusting the key. If any entry fails any check, configured_federation_peers logs the error and returns an empty list, which disables federation for the whole process. A config with one malformed peer among twenty is not partially trusted. It is not trusted.
pub struct FederationEnvelope {
pub id: FederationMessageId,
pub peer_id: FederationPeerId,
pub tenant_id: TenantId,
pub continuity_id: ContinuityId,
pub epoch: ExecutionEpoch,
pub destination_runtime_id: RuntimeId,
pub payload_sha256: String,
pub issued_at: DateTime<Utc>,
pub expires_at: DateTime<Utc>, // at most issued_at + 300 s
pub signature: String, // Ed25519 over the other fields
}POST /continuity/federation/sign builds one of these. It checks the tenant, loads the current epoch from the execution row, confirms the destination runtime is registered and not draining, refuses a TTL above 300 seconds and a payload above 16 MiB, signs, and appends a federation_signed provenance entry. It does not send anything. Delivery is a separate, explicit step by the caller, so an endpoint that can authorise a transfer cannot be turned into a request-forgery primitive, and an operator can see exactly where egress happens.
On the receiving side, verify_federation_envelope rejects an envelope that was issued in the future, has expired, is older than 300 seconds, or claims a lifetime longer than that. It rejects a revoked peer and a tenant not on that peer's allowlist. Then it recomputes the payload digest, checks the key against the trust root, and verifies the signature. Only after all of that does accept_federation_message insert a receipt with ON CONFLICT (tenant_id, peer_id, message_id) DO NOTHING, so a replayed envelope inserts zero rows and is refused. The same call deletes up to 100 expired receipts first; their envelopes can no longer verify, so keeping them would grow the table to defend against something already impossible.
Fail closed, and what that costs
Every rule above comes from one commitment: when the engine cannot prove who owns an execution, or whether an external effect happened, it stops. An unresolved receipt blocks the preview. A missing handler blocks the accept. A stale epoch is rejected at the row. One bad peer disables federation. In every case the alternative, proceed and hope, trades a visible stall for an invisible violation, and invisible violations get found by auditors and customers rather than by me.
The cost is real. Fail-closed systems stall, and stalls need people. You need a list of parked executions, tooling to revoke or retry a handoff, a reconciliation path for Unknown receipts, and an alert on the length of that queue. Without those, silent corruption becomes silent paralysis. That is better, but only somewhat.
And to repeat the caveat from the top: the pieces here are tested individually and as a server-to-device roundtrip, but I have not yet run a handoff between two organisations' deployments under real load with real failures. The parts I am least sure about are operational, not cryptographic. How long does a parked Transferring execution sit before someone notices, and how often does a device come back after a capsule has expired. Those numbers will come from running it, and I will write them up when I have them.
What each mechanism prevents
| Mechanism | Failure it prevents | Failure mode if omitted |
|---|---|---|
| Monotonic epoch, checked increment | Stale owner acting after handoff | Two runtimes dispatch the same effect |
| Transferring state | Ambiguous ownership mid-handoff | A crashed handoff leaves two owners |
| Unique index on (tenant, current_instance_id) | One local run in two executions | Effect receipts split across scopes |
| Preview digest on handoff creation | Acting on evidence nobody checked | Handoff proceeds with unresolved effects |
| Capsule AAD binding | Payload replayed elsewhere | Ciphertext re-attached to another manifest |
| Capsule schema rule | Silent field loss on transfer | A constraint disappears unnoticed |
| Capsule requirements | Landing on an incapable runtime | Execution stranded, cannot be recalled |
| Hash-chained provenance | Undetected record edits | History becomes unverifiable |
| External head anchor | Undetected chain truncation | Deleted entries verify as consistent |
| Trusted-key registry | History invalidated by rotation | Every old signature fails after a key change |
| Peer allowlist and trust root | Rogue destination | Execution handed to an attacker |
| Envelope TTL and receipt | Replayed transfer | Duplicate delivery of the same move |
Each row is cheap to build and expensive to discover missing after the fact. That asymmetry is the whole argument for the list.
Related Matching Engine Guides
The effect receipts that have to survive a relocation intact.
Why a state snapshot is what makes a capsule possible at all.
The tenant boundary that every transfer is checked against.
Claim-and-lease semantics inside one deployment, before they cross machines.
Related reading
The pillar guide covers the execution model, crate architecture, storage design, and mobile runtime that make portable execution possible.
Read the durable workflow engine architecture guide →