Blast-Radius Containment in a Multi-Tenant Engine

Orch8 runs many customers through one engine, so one customer with a broken vendor can slow everyone down. This is how I contain that at four points in the step path: the claim query, the circuit breaker, the cross-workflow handlers, and storage routing. Each one has a detail I got wrong the first time.

By Oleksii Vasylenko, Technical Lead · Published · Updated · 11 min read

Where this comes from. Orch8 is a durable workflow engine I built solo in Rust: every crate, every SDK, every integration. Ten-crate workspace, PostgreSQL and SQLite backends, server and on-device execution. The details below are the actual implementation, not a reference architecture.

Multi-tenant isolation in a workflow engine comes down to one question asked over and over: for each shared resource, where is the tenant boundary, and what happens when one tenant uses all of it? This article walks through the four places I answer that question in Orch8, a durable workflow engine I wrote in Rust, and the mistakes I made at each one.

The failure looks like this. A tenant points a workflow at an endpoint that started returning 500s at 2pm. Their steps fail and retry. Every retry takes a claim slot, a connection from the pool, and a worker task. Other tenants see their workflows slow down, then queue, then miss deadlines. Nobody on the platform did anything wrong. One customer had a vendor with a bad afternoon and the engine spread it around.

How one tenant's failing dependency degrades every other tenantFailing steps retry, retries consume claim slots, connections, and worker capacity, and unrelated tenants slow down, queue, and miss deadlines.Tenant A endpointstarts returning 500sA's steps fail and retryretries consume claim slots,connections, worker capacityTenant B slowsTenant C queuesTenant D misses deadlines
One tenant retrying against a dead endpoint eats claim slots and worker capacity. The other tenants only see the queue getting longer.

Shared infrastructure does this by default. The engine has to decide, per resource, whether the limit is per tenant or global, and then enforce it at the point where the resource is actually consumed. Wrappers and middleware get bypassed by the next code path someone adds.

A quick model of the engine, since the rest of the article uses its vocabulary. A workflow definition is a sequence of blocks. Each block names a handler, either a built-in like http_request or llm_call, or a handler the user registered. A run of a sequence is an instance: one row in task_instances with a state, a JSONB context, and a next_fire_at. The engine is one Rust binary running a tick loop against Postgres (or SQLite for tests and on-device use). Every tick claims a batch of due instances and executes their ready steps.

For one step of one instance, the tenant checks happen in this order:

  1. Claim. scheduler::process_tick calls claim_due_instances(now, limit, max_per_tenant). The SQL locks due rows with FOR UPDATE SKIP LOCKED and, if a cap is configured, keeps at most max_per_tenant rows per tenant. This is where a huge tenant is stopped from filling the batch.
  2. Pre-flight. step_exec::execute_step_block runs deadline, budget, delay, send-window, and rate-limit checks, then step_block::breaker_preflight_at. That function looks up the (tenant, handler) breaker and returns either a step definition to dispatch (possibly with the handler swapped for a fallback) or a time to defer until.
  3. Dispatch. If the handler is registered in-process it runs here; otherwise the step goes to the external worker queue. Handlers that reach other instances (emit_event, send_signal, query_instance) do their own tenant checks inside the handler body.
  4. Outcome. record_success or record_failure on the same (tenant, handler) breaker. Both retryable and permanent failures count toward tripping it.
  5. Storage. Every storage call above goes to the backend the process was booted with. The TenantPartitionRouter in orch8-storage selects a backend per tenant from a placement table, and I cover its rules below, but I should say up front that the server does not route requests through it yet. It is built and tested at the storage layer and not wired into the request path.

The rest of this article takes those in order, spending most of the time on the breaker because that is where the isolation work has the most surprising details.

A circuit breaker stops calling a dependency that is clearly down. Count consecutive failures, trip open past a threshold, reject fast during a cooldown, then let one probe through. The defaults in orch8-server/src/main.rs are 5 failures and a 60 second cooldown:

let cb_registry = Arc::new(
    CircuitBreakerRegistry::new(5, 60).with_storage(storage.clone()),
);
cb_registry.load_from_storage().await;
orch8-server/src/main.rs. The registry is built once at boot with storage injected.

What decides whether the breaker helps or hurts in a shared engine is the key. My first version keyed it by handler name. Tenant A pointed http_request at a broken endpoint, the breaker tripped, and every tenant using http_request (which is nearly all of them) got rejected for a minute. The reliability feature had turned one customer outage into a platform outage.

#[derive(Clone, Debug)]
struct Key(TenantId, String);

pub struct CircuitBreakerRegistry {
    breakers: DashMap<Key, CircuitBreakerState>,
    default_threshold: u32,
    default_cooldown_secs: u64,
    storage: Option<Arc<dyn StorageBackend>>,
    tracker: TaskTracker,
}
orch8-engine/src/circuit_breaker.rs. The tenant is part of the identity.
Circuit breaker keyed by handler alone versus by tenant and handlerKeying by handler name means one tenant's broken endpoint rejects every tenant using that handler. Keying by the tenant and handler pair contains the failure to its origin.Keyed by (tenant, handler)tenant A: http_request failsbreaker(A, http_request) openstenants B, C, D unaffectedKeyed by handler alonetenant A: http_request failsbreaker(http_request) opensEVERY tenant using http_requestis now rejected
Keyed by handler alone, one tenant trips the breaker for every tenant sharing that handler. Keyed by (tenant, handler), the trip stays inside the tenant that caused it.

The composite key creates a Rust problem. check runs on every step dispatch, and a naive DashMap<(TenantId, String), _> lookup means building an owned tuple, which means a String allocation, per check. The fix is the Borrow trick you use to look up a HashMap<String, _> with a &str, extended to a two-field key. A KeyRef(&TenantId, &str) struct and the owned Key both implement a small CircuitKey trait, Key implements Borrow<dyn CircuitKey>, and Hash and Eq are defined on the trait object so the borrowed and owned forms hash and compare identically. That last part is the requirement Borrow imposes and the one people miss. If the two forms disagree on hashing, lookups silently fail.

struct KeyRef<'a>(&'a TenantId, &'a str);

impl Hash for dyn CircuitKey + '_ {
    fn hash<H: Hasher>(&self, state: &mut H) {
        self.tenant_id().hash(state);
        self.handler().hash(state);
    }
}

pub fn check(&self, tenant_id: &TenantId, handler: &str) -> Result<(), u64> {
    let search = KeyRef(tenant_id, handler);
    let q: &dyn CircuitKey = &search;
    if let Some(breaker) = self.breakers.get(q)
        && matches!(breaker.state, BreakerState::Closed | BreakerState::HalfOpen)
    {
        return Ok(());          // read lock only, no write, no allocation
    }
    // ... Open: check cooldown under get_mut; unknown key: insert default
}
orch8-engine/src/circuit_breaker.rs. Lookups from borrowed parts, no allocation on the healthy path.

check is also careful about lock shape. A healthy breaker is answered under a DashMap read lock. Only an Open breaker takes the shard write lock, to move it to HalfOpen when the cooldown has elapsed. record_success returns early when the breaker is already Closed with zero failures, and record_failure returns early when it is already Open. Under load, nearly every call takes the read path.

An in-memory breaker resets on restart. Picture a fleet restart in the middle of a dependency outage: every tripped breaker comes back Closed, and every node resumes hammering the thing that is down, at the moment it can least cope. So Open transitions are mirrored to a circuit_breakers table (migration 029) and reloaded at boot with their opened_at, which preserves the cooldown clock.

CREATE TABLE IF NOT EXISTS circuit_breakers (
    tenant_id         TEXT        NOT NULL,
    handler           TEXT        NOT NULL,
    state             TEXT        NOT NULL,
    failure_count     INTEGER     NOT NULL,
    failure_threshold INTEGER     NOT NULL,
    cooldown_secs     BIGINT      NOT NULL,
    opened_at         TIMESTAMPTZ,
    PRIMARY KEY (tenant_id, handler)
);

CREATE INDEX IF NOT EXISTS idx_circuit_breakers_open
    ON circuit_breakers (state) WHERE state = 'open';
migrations/029_circuit_breakers.sql. Primary key is the same (tenant, handler) pair.

Only Open rows get written. The reasons:

  • Closed is the default. A row for every untripped breaker would mean writing a record to say nothing is wrong, for every (tenant, handler) pair that ever ran.
  • HalfOpen is a probe owned by the live process. After a restart there is no process running that probe, so the state has no meaning. A persisted Open row whose cooldown has already elapsed is flipped to HalfOpen by the first check after boot, which gives the same effect without storing it.
  • Open is the one state whose loss changes behaviour for the worse. When the breaker moves from Open to HalfOpen or Closed, the row is deleted so a crash cannot revive a breaker that had already recovered.
Circuit breaker states and which one is persistedClosed, Open, and HalfOpen. Only Open is written to storage and rehydrated at boot with its cooldown intact, because it is the only state whose loss on restart changes behaviour badly.failures ≥ thresholdcooldown elapsedprobe succeedsprobe failsClosedOpenHalfOpenthe ONLY state writtento storage and rehydratedat boot, cooldown intactdefault — persisting itwould write a row forevery untripped breaker
Only the Open transition writes a row. Closed is implied by absence, and HalfOpen belongs to a process that will not exist after a restart.

The write is off the hot path. A transition spawns the upsert or delete onto a tokio_util::task::TaskTracker, so check, record_failure, record_success, and reset stay synchronous. The tracker exists because the first version used a bare tokio::spawn, and a write in flight during runtime shutdown could be aborted halfway, leaving the table disagreeing with what the process had in memory. Now the server calls flush at shutdown, which closes the tracker and awaits every outstanding write.

The breaker exists for external dependencies: HTTP APIs, LLM providers, plugin sidecars, gRPC services. Put one on a handler with no external dependency and you only get false positives. A fail step in a test suite would trip the breaker for fail. A send_signal that loses a race with the target completing would count as a failure. Neither should take down every instance using that handler for a minute.

pub fn is_breaker_tracked(handler: &str) -> bool {
    !matches!(
        handler,
        "noop" | "log" | "sleep" | "fail" | "self_modify"
            | "emit_event" | "send_signal" | "query_instance" | "human_review"
    )
}
orch8-engine/src/circuit_breaker.rs. One list, next to the breaker, rather than a flag on each handler.

I keep the skip-list here instead of as a property on each handler because it is policy that has to be reviewed as a set. When the list lives in one place, "which handlers bypass the breaker?" is a five-second answer. Scattered opt-outs mean grepping, and a wrong entry hides for a long time.

When a breaker is open the engine does not simply fail the step. breaker_preflight_at is a pure decision function with no storage writes. If the step declares a fallback_handler and that handler's own breaker is not open (or it is on the skip-list), the step definition is cloned with the handler swapped and dispatch continues. Otherwise the function returns a fire_at equal to now plus the remaining cooldown, and the caller moves the instance from Running back to Scheduled at that time. The decision is shared by both dispatch paths, the flat step loop in step_exec and the tree evaluator in step_block, so open-breaker behaviour cannot drift between them.

pub(crate) enum BreakerDecision<'a> {
    Proceed(Cow<'a, StepDef>),
    Defer { fire_at: DateTime<Utc> },
}

let Some(fb) = step_def.fallback_handler.as_deref() else {
    let fire_at = now + Duration::seconds(remaining_secs as i64);
    return BreakerDecision::Defer { fire_at };
};
// re-check the fallback's own breaker, then:
let mut cloned = step_def.clone();
cloned.handler = fb.to_string();
BreakerDecision::Proceed(Cow::Owned(cloned))
orch8-engine/src/handlers/step_block.rs. Proceed with a possibly-swapped step, or defer.

Breakers handle a tenant whose dependency is failing. They do nothing about a tenant who is simply large. If the claim query takes rows in priority-then-time order, a tenant with 500,000 due instances fills every batch, and everyone else waits with no error anywhere. Just a queue that never reaches them.

The Postgres claim query caps rows per tenant using ROW_NUMBER() OVER (PARTITION BY tenant_id). Postgres will not allow FOR UPDATE SKIP LOCKED in the same query level as a window function, so the rows are locked in a CTE first and ranked in the outer query. The CTE over-selects limit * max_per_tenant rows, because if it locked only limit rows and one tenant owned all of them, the cap would trim the batch to almost nothing. The full story of that query, including the version that was wrong, is in the scheduler article linked below.

WITH locked AS (
    SELECT id, tenant_id, priority, next_fire_at
    FROM task_instances
    WHERE (next_fire_at IS NULL OR next_fire_at <= $1) AND state = 'scheduled'
    ORDER BY priority DESC, next_fire_at ASC NULLS FIRST
    LIMIT $4                       -- limit * max_per_tenant
    FOR UPDATE SKIP LOCKED
), ranked AS (
    SELECT id, tenant_id, priority, next_fire_at,
           ROW_NUMBER() OVER (PARTITION BY tenant_id
                              ORDER BY priority DESC, next_fire_at ASC NULLS FIRST) AS rn
    FROM locked
), winners AS (
    SELECT id FROM ranked WHERE rn <= $3
    ORDER BY priority DESC, next_fire_at ASC NULLS FIRST
    LIMIT $2
)
SELECT task_instances.* FROM task_instances JOIN winners USING (id)
orch8-storage/src/postgres/instances.rs, claim_due. Lock, rank, trim, then load the winners.

One thing I should be clear about: the cap is opt-in. ORCH8_MAX_INSTANCES_PER_TENANT defaults to 0, which means no limit, and the query falls back to a plain LIMIT $2 FOR UPDATE SKIP LOCKED. A single-tenant deployment pays nothing for fairness it does not need. A multi-tenant deployment has to set it, and I would rather that be a visible configuration decision than a default that changes claim behaviour for everyone.

An engine that lets workflows spawn, signal, and query each other has created a way for one tenant to reach another. The enforcement has to be inside the handler, where the target instance is loaded, and not in a wrapper that a future code path can skip.

  • emit_event spawns a child instance in the caller's tenant only. An optional dedupe_key can be scoped to the parent instance or to the whole tenant, and the dedupe row plus the child instance are written in one storage transaction via create_instance_with_dedupe, so a crash mid-operation cannot leave a dedupe row pointing at a child that was never created.
  • send_signal loads the target, checks target.tenant_id == ctx.tenant_id, and only then calls enqueue_signal_if_active. That method re-reads the target state under a row lock (FOR UPDATE on Postgres, write-transaction semantics on SQLite) and refuses terminal targets in the same transaction as the insert. The first version checked state, then inserted, and a target could complete in between.
  • query_instance reads same-tenant instances only. A missing target and a target in another tenant both return { found: false }.
let target = match target {
    Some(t) if t.tenant_id == ctx.tenant_id => t,
    _ => return Ok(json!({ "found": false })),
};
orch8-engine/src/handlers/query_instance.rs. Absent and foreign are the same answer.

That last one is the kind of leak that gets through a review. The access check is correct and no data crosses the boundary, but if "does not exist" and "exists but is not yours" produce different responses, the response itself tells an attacker whether an instance ID is live in another tenant. send_signal does the same thing with its error: cross-tenant and not-found both return the message "target instance not found", and there is a test asserting the cross-tenant error does not mention the word terminal, because that would reveal the target exists and has finished.

Logical isolation has a floor. Some tenants need their rows in a specific country, or on hardware their contract names, or simply not next to anyone else. That means routing a tenant to its own storage backend, and routing is the place where a convenient default turns into a data breach.

The TenantPartitionRouter in orch8-storage/src/tenant_partition.rs reads a placement record for every route and fails closed. A tenant with no record gets NotFound. A record naming a backend that this process did not register at boot gets Unsupported. Neither falls back to another partition. "We could not find your placement so we used the default one" is how a regulated tenant's data ends up in the wrong jurisdiction, and I would rather the request fail.

pub async fn route(&self, tenant_id: &TenantId) -> Result<RoutedTenantStorage, StorageError> {
    let placement = self.placements.get_tenant_placement(tenant_id).await?
        .ok_or_else(|| StorageError::NotFound {
            entity: "tenant storage placement", id: tenant_id.to_string(),
        })?;
    let backend = self.backends.get(&placement.backend_id).ok_or_else(|| {
        StorageError::Unsupported(format!(
            "tenant '{}' is assigned to unregistered storage backend '{}'",
            tenant_id, placement.backend_id))
    })?;
    Ok(RoutedTenantStorage { placement, backend: Arc::clone(backend) })
}
orch8-storage/src/tenant_partition.rs. Two error paths, no fallback branch.
Fail-closed tenant storage partition routingA missing placement record returns NotFound and an unregistered backend returns Unsupported. Neither falls back, because there is deliberately no default backend.noyesnoyestenant-scoped requestparse and validate TenantIdplacement recordexists?NotFound — fail closedbackend registeredin this process?Unsupported — fail closedrun the WHOLE operationon that backendthere is nodefault backend
Both error paths terminate the request. There is no default backend for either of them to fall through to.

Placement changes are epoch-fenced, the same way ownership is fenced everywhere else in the engine. The table (migration 078) has tenant_id as primary key, a backend_id, and a positive epoch. advance_tenant_placement is a conditional upsert: insert if absent, otherwise replace only when the proposed epoch is strictly greater than the stored one. A stale control-plane writer replaying an old placement gets a Conflict and changes nothing.

INSERT INTO tenant_storage_placements (tenant_id, backend_id, epoch, updated_at)
VALUES ($1, $2, $3, $4)
ON CONFLICT (tenant_id) DO UPDATE
    SET backend_id = EXCLUDED.backend_id,
        epoch      = EXCLUDED.epoch,
        updated_at = EXCLUDED.updated_at
    WHERE EXCLUDED.epoch > tenant_storage_placements.epoch
-- rows_affected != 1  =>  StorageError::Conflict
orch8-storage/src/tenant_partition.rs (Postgres). Older epochs match zero rows.

backend_id is an operator-chosen name, validated to 1-128 characters of letters, digits, dots, dashes, and underscores. It is never a connection string. The placement table belongs in a highly available control-plane store and may be replicated more widely than credentials should travel. Connection details stay in each process's protected configuration, so the routing table is authoritative without being secret.

The router selects a backend. It does not copy data, coordinate dual writes, or make a multi-backend transaction atomic. Moving a tenant means quiescing writes, copying, validating, registering the destination on every serving node, and only then advancing the epoch. And as I said earlier, this component exists in the storage crate with its own tests and doc, and the server still boots against one backend. Wiring the router into the request path is the next piece of work, not finished work.

None of this is specific to a workflow engine. Any shared service has the same list of resources and the same question per resource.

  1. Write down the shared resources: connection pools, worker capacity, rate limits, caches, breakers, queues, scheduler batches, memory.
  2. For each one, decide whether the boundary is per tenant, global, or tiered. An undecided boundary is a global one.
  3. Put cross-tenant checks inside the operation that loads the foreign object. Middleware gets bypassed.
  4. Return the same response for "does not exist" and "not yours". Different errors are an existence oracle.
  5. Fail closed on routing. A default backend is convenient right up until the first compliance incident.
  6. Run the noisy-neighbour test on purpose. One tenant at ten times everyone else's load, and assert the others still hit their objectives. I did not do this early enough, and the handler-keyed breaker would have shown up on day one if I had.

The reason to do this early is that isolation is hard to retrofit. Every boundary has to exist at the point where the resource is consumed. With five call sites, adding a tenant dimension is an afternoon. With fifty, it is a migration.

MechanismBoundaryContainsCost
Per-tenant circuit breaker(tenant, handler)One tenant's broken dependencyMore breaker state in memory; a Borrow impl to avoid allocating per check
Persisted Open statecircuit_breakers tableFleet restart resuming a stampedeSmall crash window before the tracked write lands
Breaker skip-listHandler classFalse positives on internal stepsA list to keep correct
Fallback handler on open breakerStep definitionA stalled step when an alternative existsPer-step configuration
Per-tenant claim capScheduler batchOne tenant monopolising throughputOver-select in the claim CTE; opt-in via config
Tenant-scoped coordinationInside emit_event, send_signal, query_instanceCross-tenant spawns, signals, and readsChecks on every coordination path
Uniform not-found responseHandler outputCross-tenant existence oracleLess specific error messages
Epoch-fenced placementtenant_storage_placementsStale routing after a moveExplicit migration procedure
No default backendTenantPartitionRouterData landing in the wrong partitionEvery tenant needs a record; not yet wired into the server

None of these is expensive on its own. The expensive version is adding them to a system that assumed one shared pool for two years.

The pillar guide covers the full engine architecture: execution model, crate design, storage backends, and the operational decisions behind them.

Read the durable workflow engine architecture guide →