Moving a Running Workflow Across a Trust Boundary

Copying a stopped process is a file copy. Moving a running one means transferring authority over its side effects. This is how Orch8 does that: an ownership row with an epoch, a signed and encrypted capsule, effect receipts keyed to the logical execution, and a provenance chain that anyone can verify.

By Oleksii Vasylenko, Technical Lead · Published · Updated · 16 min read

Where this comes from, and what state it is in. Orch8 is a durable workflow engine I built solo in Rust. The ownership model, capsule format, effect receipts, provenance chain, and federation envelope described here are real code with unit and integration tests, and the server-to-device roundtrip runs in the mobile test suite. What has not happened yet is a handoff across a live trust boundary between two organisations at production scale. Read this as a design I can defend line by line, not as a system with production incidents behind it.

Portable workflow execution means taking a workflow that is partway through its steps and continuing it on a different machine, possibly one owned by someone else. Three situations pushed me into building it, and restarting the workflow from scratch on the new machine does not solve any of them.

  • A phone goes offline in the middle of an onboarding flow the server was driving. The remaining steps should run on the device and reconcile when the network comes back.
  • One step has to touch data that cannot leave a customer's own infrastructure or a jurisdiction. Only that step needs to run there. The rest of the orchestration does not.
  • A process spans two companies. Neither will run the other's code or open its database, but the process is one process and needs one continuous record.

My first attempt was the obvious one: serialise the instance, send it, deserialise it, carry on. It works until the source does not actually stop. Then two runtimes both think they own the execution, and both call the payment API. Everything in this article exists to make that state unreachable, and most of it is checked by the database rather than by application code.

Some vocabulary first, because the rest of the article uses it. In Orch8 a workflow definition is a sequence made of blocks. Each block names a handler, either a built-in such as http_request or one a user registered. A run of a sequence is an instance: one row in task_instances with a state, a JSONB context, and a next_fire_at timestamp. The engine is one Rust binary that runs a tick loop against Postgres or SQLite, claims due instances, and executes their next block. That is the whole engine for a workflow that never leaves its server.

Portable continuity adds a second identity on top of the instance. An instance ID is local to one runtime and changes every time the execution moves. The thing that stays constant is the continuity_id, and its authoritative record is one row in continuity_executions:

pub struct ContinuityExecution {
    pub continuity_id: ContinuityId,     // stable across every move
    pub tenant_id: TenantId,
    pub current_instance_id: InstanceId, // the local run on whichever runtime owns it now
    pub owner_runtime_id: RuntimeId,     // who may act right now
    pub epoch: ExecutionEpoch,           // u64, increments on every ownership change
    pub state: OwnershipState,           // Owned | Transferring | Completed
    pub updated_at: DateTime<Utc>,
}

CREATE TABLE continuity_executions (
    continuity_id    UUID PRIMARY KEY,
    tenant_id        TEXT NOT NULL,
    epoch            BIGINT NOT NULL CHECK (epoch >= 0),
    owner_runtime_id UUID NOT NULL,
    state            TEXT NOT NULL,
    record           JSONB NOT NULL,
    updated_at       TIMESTAMPTZ NOT NULL,
    UNIQUE (tenant_id, continuity_id)
);
orch8-types/src/continuity.rs and migrations/057_portable_continuity.sql. The epoch column is the fencing token; the state column is what makes a half-finished handoff safe.

Alongside it sit five more tables from the same migration: execution_handoffs (one row per transfer attempt, with its own version counter), execution_capsules (signed manifests with an expiry), runtime_capabilities (what each registered runtime says it can do, with an expiry), effect_receipts (every external side effect, keyed to the continuity ID rather than the instance), and provenance_entries (the hash chain). A later migration, 061, adds a unique index on (tenant_id, current_instance_id), so a local instance can belong to at most one continuity execution. I would rather have that as an index than as a check in Rust that a future code path can skip.

The API surface is in orch8-api/src/continuity.rs and the transition rules in orch8-engine/src/continuity.rs. A server-to-server handoff inside one deployment goes through six calls. Every one of them is a compare-and-swap against a row it read a moment earlier.

  1. Preview. POST /continuity/executions/{id}/handoff/preview names a destination runtime and the capsule requirements. build_handoff_preview loads the destination's live capabilities, runs assess_compatibility, runs the placement policy, and lists every effect receipt that is not resolved. It returns a preview_sha256 over all of that evidence. If any effect receipt is still Dispatched or Unknown, the preview reports compatible: false. You cannot move an execution while it has a side effect in flight whose outcome nobody knows.
  2. Create. POST .../handoff must carry that digest. create_handoff rebuilds the preview and refuses if the digest no longer matches, so a handoff can only be created against evidence the caller actually looked at. The handoff row starts in Requested with version = 0.
  3. Export. export_handoff checks that the source instance is Paused or Waiting, moves the handoff to Quiescing with a CAS on its version, builds the capsule (next section), and then calls commit_handoff_export. That single storage transaction flips the handoff to Exported and the execution from Owned to Transferring, with WHERE epoch = $expected AND owner_runtime_id = $expected AND state = 'owned' on the update. If zero rows change, the whole thing rolls back and the caller gets StaleClaim.
  4. Import. The destination calls import_capsule with the signed manifest and the encrypted payload. verify_and_import_paused_capsule_bytes checks the manifest signature against the engine's key, the payload length and SHA-256 against the manifest, decrypts, validates bounds, and creates a new local instance in Paused. Nothing about ownership changes yet.
  5. Accept. accept_handoff computes epoch.checked_next() and writes the new owner. The storage call updates continuity_executions only where the epoch, owner, and state = 'transferring' still match, and only if is_capsule_import_instance confirms the destination instance was created from this handoff's capsule. Then it appends a runtime_claimed provenance entry.
  6. Resume and complete. resume_handoff moves the destination instance to Scheduled so the tick loop picks it up, and the handoff row walks Accepted to Resumed to Completed. The HandoffState::can_transition_to table has no path from Requested to Resumed, and no path out of a terminal state, so a client cannot skip the middle.

The mobile path is the same protocol with fewer round trips. orch8-mobile/src/continuity.rs exposes export_capsule, import_capsule, and activate_capsule. The device signs the manifest with its own Ed25519 key (the server learns that public key at registration through capsule_signing_public_key), and activate_capsule does the ownership CAS locally with cas_continuity_owner. The test portable_roundtrip_survives_redelivery_and_activation runs an export, imports the same capsule twice, activates it, and checks that the second import is a no-op rather than a second owner.

A boolean is_owner cannot survive a network partition. The old owner still has its flag set and no way to learn that the world moved on. What works is a generation number that increases on every transfer, so that a write from a previous owner is recognisably from the past rather than merely late. This is the fencing-token idea from Martin Kleppmann's post on distributed locking, applied to execution authority instead of a lock.

pub struct ExecutionEpoch(u64);

impl ExecutionEpoch {
    pub fn checked_next(self) -> Result<Self, ContinuityError> {
        self.0
            .checked_add(1)
            .map(Self)
            .ok_or(ContinuityError::EpochOverflow)
    }
}
orch8-types/src/continuity.rs. The increment is checked. A wrapped epoch would re-authorise every stale owner in the history of the execution.
Handing a running execution from server to deviceThe record moves to Transferring so neither side may dispatch effects, the destination verifies it can run the capsule, ownership is claimed at the next epoch, and a late write from the old owner is rejected as stale.DeviceContinuity recordServer (epoch 4)neither side may dispatcheffects in this windowrejected as Stale,not retriedstate = Transferring1capsule + requirements2verify handlers, credentials,region, hardware3claim ownership, epoch 4 → 54Owned at epoch 55late write at epoch 46
A late write from the old owner carries epoch N after the row has moved to N+1. The storage layer rejects it in the WHERE clause, so the application never has to decide whether the old owner is slow or dead.

The Transferring state is the second half of this. Between the export commit and the accept, neither runtime may dispatch effects: the source has given up Owned, and the destination has not yet claimed it. If the process dies in that window, the execution is parked in a state that says so. An operator can revoke the handoff (Exported to Revoked is a legal transition) and the source reclaims it at a fresh epoch. Without that intermediate state the failure mode is two rows of evidence that both say "mine".

Execution ownership statesOwned, Transferring, and Completed. The Transferring state creates a window in which neither runtime may dispatch effects, so a crashed handoff parks the execution rather than leaving two owners.handoff beginsclaimed at epoch+1transfer abortedexecution finishedOwnedTransferringCompletedno effects may bedispatched by either side
Owned, Transferring, Completed. A crash during Transferring leaves the execution parked with no owner, which is recoverable. Two owners is not.

The capsule is what actually crosses the wire. It has two parts: a signed manifest that travels as JSON, and an encrypted payload that is stored as an artifact and fetched or attached separately. export_paused_capsule_manifest in orch8-engine/src/capsule.rs assembles both.

pub struct CapsulePayload {
    pub instance: CapsuleInstanceState,   // sequence_id, namespace, priority, timezone,
                                          // metadata, context, budget, parent_instance_id
    pub checkpoint: Checkpoint,           // the latest durable checkpoint
    pub outputs: Vec<BlockOutput>,        // every completed block's output
    pub pending_waits: Vec<Value>,
    pub pending_signals: Vec<Value>,
    pub effect_ids: Vec<EffectId>,        // receipts that belong to this execution
    pub artifacts: Vec<ArtifactReference>,
    pub stream_cursors: Vec<StreamCursor>,
    pub redacted_audit_context: Value,
}
// MAX_OUTPUTS 10_000, MAX_PENDING_ITEMS 10_000, MAX_ENCODED_BYTES 64 MiB

pub struct CapsuleManifest {
    pub schema: CapsuleSchemaVersion,     // { major, minor }
    pub capsule_id: CapsuleId,
    pub continuity_id: ContinuityId,
    pub source_instance_id: InstanceId,
    pub epoch: ExecutionEpoch,
    pub tenant_id: TenantId,
    pub source_runtime_id: RuntimeId,
    pub allowed_destination_runtime_id: Option<RuntimeId>,
    pub sequence: SequenceIdentity,       // id, version, content_sha256
    pub checkpoint: CheckpointIdentity,   // block_id, sha256
    pub requirements_sha256: String,
    pub payload_artifact: ArtifactReference, // key, sha256, bytes
    pub provenance_head: Option<String>,
    pub issued_at: DateTime<Utc>,
    pub expires_at: DateTime<Utc>,        // default 300 s, at most 3600 s
    pub signing_key_id: String,
    pub encryption_key_id: String,
}
orch8-types/src/continuity.rs. The payload is a deliberate subset of the instance row, so the capsule format does not track the engine's internal SQL schema.

The payload is canonical JSON, encrypted with the engine's field encryptor using additional authenticated data made of the tenant, the continuity ID, the epoch, and the destination runtime. So a payload cannot be re-attached to a different manifest, replayed into a different tenant, or delivered to a runtime other than the one it was exported for, even by someone who has the ciphertext. For a device transfer the caller can supply a one-time transfer key instead of the engine key, which is what the mobile path does.

Export refuses unless the source instance is Paused or Waiting. A block that is mid-execution has no clean boundary to cut at, and the checkpoint the manifest names would not describe the state that actually left. The manifest also records the sequence content hash and the provenance head at export time, so the destination can tell whether it is looking at the same workflow definition and the same history the source had.

This is where the naive copy fails hardest. Say the workflow charged a card on the server and then relocated to a phone. If the record of that charge does not travel, or travels but is not the record the phone consults, the phone re-runs the step. The at-most-once design in the effect receipts article only works if the receipt is keyed to something that survives relocation.

So an EffectReceipt carries continuity_id and epoch, and only secondarily instance_id. The receipts table has a foreign key to continuity_executions, not to the instance. When the execution lands on a new runtime with a new local instance, the receipts still belong to the same logical execution, and the capsule payload lists their IDs so the destination knows which ones to expect.

match storage.dispatch_effect_receipt_at_most_once(&tenant, &updated).await? {
    EffectDispatchOutcome::Dispatched => { /* this caller owns the effect */ }
    EffectDispatchOutcome::Duplicate  => Err(EngineError::InvariantViolation { .. }),
    EffectDispatchOutcome::Stale      => Err(blocked(&current)),   // older epoch tried to act
}
orch8-engine/src/effect_guard.rs. The storage layer returns three different outcomes for a dispatch attempt, and the epoch is what separates the last two.

Stale and Duplicate are separate on purpose. A duplicate means the same execution tried to dispatch one effect twice, which is a workflow or engine bug. A stale outcome means a runtime from a superseded epoch tried to act after ownership moved, which is a handoff race. The two need different responses and different people looking at them.

And, as mentioned in the walk-through, the handoff preview refuses to report a destination as compatible while any receipt is unresolved. A Dispatched receipt with no answer yet, or an Unknown one waiting for reconciliation, blocks the move. I would rather delay a transfer than transfer an execution whose recent past nobody can vouch for.

A workflow that needs an S3 credential and a GPU cannot run on a phone. Finding that out after accept means an execution stranded on a runtime that can never advance it, and since ownership already moved, the origin cannot take it back without another handoff. So the requirements travel with the request and are checked before anything changes.

pub struct CapsuleRequirements {
    pub handlers: Vec<String>,        // every block handler must be registered there
    pub plugins: Vec<String>,
    pub credentials: Vec<String>,
    pub regions: Vec<String>,         // residency: destination must be in one of them
    pub hardware: Vec<String>,
    pub requires_network: bool,
    pub requires_human_ui: bool,      // only Mobile, Desktop, Browser can satisfy this
    pub minimum_trust: Option<RuntimeTrustLevel>,  // Unverified < Registered < Signed < Attested
}

pub enum RuntimeKind { Server, Edge, Mobile, Desktop, Browser }
orch8-types/src/continuity.rs. Credentials are names of configured bindings, never the secrets.

Runtimes register their capabilities with an expiry of at most five minutes, so a stale advertisement ages out on its own. Two functions read those facts. assess_compatibility produces findings with codes like HANDLERS_MISSING, REGION_NOT_ALLOWED, and TRUST_TOO_LOW, and it distinguishes Fail from Unknown: a runtime that did not report its connectivity gets NETWORK_UNKNOWN, which shows up in the preview. CapsuleRequirements::is_satisfied_by is the one used when a runtime actually claims work, and it treats unknown the same as failed. A preview can tell you "we are not sure". A claim cannot proceed on "not sure".

Delegating a slice of work to a device adds more checks on top: the current ownership epoch, live same-tenant registrations for both source and destination, every handler the isolated sub-sequence needs, a signed continuation grant bound to that destination, and a one-time token whose SHA-256 is stored and consumed on use. The device gets an isolated sub-sequence, not the whole execution, so a compromised phone can only reach its own slice.

Once an execution has crossed three runtimes and two organisations, "what happened?" needs an answer that does not depend on trusting whoever is answering. Each boundary (capsule exported, runtime claimed, effect resolved, federation envelope signed) appends a ProvenanceEntry whose hash commits to the previous entry's hash.

pub struct ProvenanceEntry {
    pub continuity_id: ContinuityId,
    pub epoch: ExecutionEpoch,
    pub kind: String,                      // "capsule_exported", "runtime_claimed", ...
    pub redacted_summary: Option<String>,  // operator-safe, bounded
    pub payload_sha256: String,
    pub previous_sha256: Option<String>,   // the chain link
    pub entry_sha256: String,
    pub signing_key_id: Option<String>,
    pub signature: Option<String>,         // Ed25519 over entry_sha256
    pub created_at: DateTime<Utc>,
}
orch8-types/src/continuity.rs. Only digests and a bounded summary. The payload itself never enters the chain.

The entry hash is SHA-256 over a domain separator (orch8-provenance-v2), the continuity ID, the big-endian epoch, and length-framed copies of the kind, summary, payload digest, and previous hash. The length framing is there so that two different field splits cannot produce the same bytes. Storing digests instead of payloads is what makes the chain safe to retain and safe to show a counterparty: a regulator can verify that a decision was recorded and never edited without the audit log becoming a second copy of the customer data.

Hash-chained provenance and why it needs an external anchorEach entry commits to its predecessor's hash. Truncating the tail leaves a chain that still verifies internally, so detecting deletion requires an expected head retained outside the execution database.comparechain still verifiesinternallyentry 1payload_sha256prev = nullentry 2payload_sha256prev = hash(1)entry 3payload_sha256prev = hash(2)headexpected headheld OUTSIDE theexecution databasetruncate entries 2-3
Each entry commits to its predecessor. Truncating the tail leaves a shorter chain that still verifies on its own, which is why verification takes an expected head from outside the database.

Key rotation uses a registry rather than re-signing history. The active continuity signing key is trusted automatically; retired key IDs map to their public keys through ORCH8_CONTINUITY_TRUSTED_SIGNING_KEYS_JSON, and verify_provenance_chain_with_keys looks each entry's signing_key_id up there. An entry signed by a key that is not in the registry fails with PROVENANCE_SIGNING_KEY_UNTRUSTED. Without that map, the first rotation would invalidate every entry signed before it.

Federation is where the trust assumptions change. Inside one deployment every runtime is mine. Across a federation boundary the destination belongs to someone else, and the design has to assume it might be wrong or hostile.

Peers come from operator configuration only, through ORCH8_FEDERATION_PEERS. They are never discovered and never self-register. parse_federation_peers in orch8-server/src/main.rs caps the list at 128 peers with unique IDs, requires an https:// endpoint under 2,048 bytes, a tenant allowlist of 1 to 256 entries, and a 64-character hex trust_root_sha256. That digest is the SHA-256 of the peer's Ed25519 public key bytes, and the verifier recomputes it before trusting the key. If any entry fails any check, configured_federation_peers logs the error and returns an empty list, which disables federation for the whole process. A config with one malformed peer among twenty is not partially trusted. It is not trusted.

pub struct FederationEnvelope {
    pub id: FederationMessageId,
    pub peer_id: FederationPeerId,
    pub tenant_id: TenantId,
    pub continuity_id: ContinuityId,
    pub epoch: ExecutionEpoch,
    pub destination_runtime_id: RuntimeId,
    pub payload_sha256: String,
    pub issued_at: DateTime<Utc>,
    pub expires_at: DateTime<Utc>,   // at most issued_at + 300 s
    pub signature: String,           // Ed25519 over the other fields
}
orch8-types/src/continuity_advanced.rs. Everything the receiver needs to bind the message to one peer, one tenant, one execution, one epoch, and one destination.

POST /continuity/federation/sign builds one of these. It checks the tenant, loads the current epoch from the execution row, confirms the destination runtime is registered and not draining, refuses a TTL above 300 seconds and a payload above 16 MiB, signs, and appends a federation_signed provenance entry. It does not send anything. Delivery is a separate, explicit step by the caller, so an endpoint that can authorise a transfer cannot be turned into a request-forgery primitive, and an operator can see exactly where egress happens.

Signing and verifying a federated transfer envelopeThe sending engine checks tenant, epoch, and destination liveness, then signs an envelope with a five-minute maximum lifetime. Signing performs no network call. The receiver verifies with its configured key for that peer and records a one-delivery receipt.Peer B engineOperator / transportPeer A enginesigning does NOT makean outbound requestcheck tenant, epoch,destination liveness1signed envelope (Ed25519, TTL≤ 300s)2envelope + unchanged payload3verify with the CONFIGURED keyfor peer A, never a request key4durable receipt — onedelivery per(tenant, peer, message)5accepted, digest into provenance6
Signing and delivery are separate calls. The receiver verifies with the key it has on file for that peer and never with a key carried in the request.

On the receiving side, verify_federation_envelope rejects an envelope that was issued in the future, has expired, is older than 300 seconds, or claims a lifetime longer than that. It rejects a revoked peer and a tenant not on that peer's allowlist. Then it recomputes the payload digest, checks the key against the trust root, and verifies the signature. Only after all of that does accept_federation_message insert a receipt with ON CONFLICT (tenant_id, peer_id, message_id) DO NOTHING, so a replayed envelope inserts zero rows and is refused. The same call deletes up to 100 expired receipts first; their envelopes can no longer verify, so keeping them would grow the table to defend against something already impossible.

Every rule above comes from one commitment: when the engine cannot prove who owns an execution, or whether an external effect happened, it stops. An unresolved receipt blocks the preview. A missing handler blocks the accept. A stale epoch is rejected at the row. One bad peer disables federation. In every case the alternative, proceed and hope, trades a visible stall for an invisible violation, and invisible violations get found by auditors and customers rather than by me.

The cost is real. Fail-closed systems stall, and stalls need people. You need a list of parked executions, tooling to revoke or retry a handoff, a reconciliation path for Unknown receipts, and an alert on the length of that queue. Without those, silent corruption becomes silent paralysis. That is better, but only somewhat.

And to repeat the caveat from the top: the pieces here are tested individually and as a server-to-device roundtrip, but I have not yet run a handoff between two organisations' deployments under real load with real failures. The parts I am least sure about are operational, not cryptographic. How long does a parked Transferring execution sit before someone notices, and how often does a device come back after a capsule has expired. Those numbers will come from running it, and I will write them up when I have them.

MechanismFailure it preventsFailure mode if omitted
Monotonic epoch, checked incrementStale owner acting after handoffTwo runtimes dispatch the same effect
Transferring stateAmbiguous ownership mid-handoffA crashed handoff leaves two owners
Unique index on (tenant, current_instance_id)One local run in two executionsEffect receipts split across scopes
Preview digest on handoff creationActing on evidence nobody checkedHandoff proceeds with unresolved effects
Capsule AAD bindingPayload replayed elsewhereCiphertext re-attached to another manifest
Capsule schema ruleSilent field loss on transferA constraint disappears unnoticed
Capsule requirementsLanding on an incapable runtimeExecution stranded, cannot be recalled
Hash-chained provenanceUndetected record editsHistory becomes unverifiable
External head anchorUndetected chain truncationDeleted entries verify as consistent
Trusted-key registryHistory invalidated by rotationEvery old signature fails after a key change
Peer allowlist and trust rootRogue destinationExecution handed to an attacker
Envelope TTL and receiptReplayed transferDuplicate delivery of the same move

Each row is cheap to build and expensive to discover missing after the fact. That asymmetry is the whole argument for the list.

The pillar guide covers the execution model, crate architecture, storage design, and mobile runtime that make portable execution possible.

Read the durable workflow engine architecture guide →