Blast-Radius Containment in a Multi-Tenant Engine
Orch8 runs many customers through one engine, so one customer with a broken vendor can slow everyone down. This is how I contain that at four points in the step path: the claim query, the circuit breaker, the cross-workflow handlers, and storage routing. Each one has a detail I got wrong the first time.
By Oleksii Vasylenko, Technical Lead · Published · Updated · 11 min read
Where this comes from. Orch8 is a durable workflow engine I built solo in Rust: every crate, every SDK, every integration. Ten-crate workspace, PostgreSQL and SQLite backends, server and on-device execution. The details below are the actual implementation, not a reference architecture.
How one tenant slows down everyone else
Multi-tenant isolation in a workflow engine comes down to one question asked over and over: for each shared resource, where is the tenant boundary, and what happens when one tenant uses all of it? This article walks through the four places I answer that question in Orch8, a durable workflow engine I wrote in Rust, and the mistakes I made at each one.
The failure looks like this. A tenant points a workflow at an endpoint that started returning 500s at 2pm. Their steps fail and retry. Every retry takes a claim slot, a connection from the pool, and a worker task. Other tenants see their workflows slow down, then queue, then miss deadlines. Nobody on the platform did anything wrong. One customer had a vendor with a bad afternoon and the engine spread it around.
Shared infrastructure does this by default. The engine has to decide, per resource, whether the limit is per tenant or global, and then enforce it at the point where the resource is actually consumed. Wrappers and middleware get bypassed by the next code path someone adds.
Where each boundary sits in the step path
A quick model of the engine, since the rest of the article uses its vocabulary. A workflow definition is a sequence of blocks. Each block names a handler, either a built-in like http_request or llm_call, or a handler the user registered. A run of a sequence is an instance: one row in task_instances with a state, a JSONB context, and a next_fire_at. The engine is one Rust binary running a tick loop against Postgres (or SQLite for tests and on-device use). Every tick claims a batch of due instances and executes their ready steps.
For one step of one instance, the tenant checks happen in this order:
- Claim.
scheduler::process_tickcallsclaim_due_instances(now, limit, max_per_tenant). The SQL locks due rows withFOR UPDATE SKIP LOCKEDand, if a cap is configured, keeps at mostmax_per_tenantrows per tenant. This is where a huge tenant is stopped from filling the batch. - Pre-flight.
step_exec::execute_step_blockruns deadline, budget, delay, send-window, and rate-limit checks, thenstep_block::breaker_preflight_at. That function looks up the(tenant, handler)breaker and returns either a step definition to dispatch (possibly with the handler swapped for a fallback) or a time to defer until. - Dispatch. If the handler is registered in-process it runs here; otherwise the step goes to the external worker queue. Handlers that reach other instances (
emit_event,send_signal,query_instance) do their own tenant checks inside the handler body. - Outcome.
record_successorrecord_failureon the same(tenant, handler)breaker. Both retryable and permanent failures count toward tripping it. - Storage. Every storage call above goes to the backend the process was booted with. The
TenantPartitionRouterinorch8-storageselects a backend per tenant from a placement table, and I cover its rules below, but I should say up front that the server does not route requests through it yet. It is built and tested at the storage layer and not wired into the request path.
The rest of this article takes those in order, spending most of the time on the breaker because that is where the isolation work has the most surprising details.
Circuit breakers keyed by (tenant, handler)
A circuit breaker stops calling a dependency that is clearly down. Count consecutive failures, trip open past a threshold, reject fast during a cooldown, then let one probe through. The defaults in orch8-server/src/main.rs are 5 failures and a 60 second cooldown:
let cb_registry = Arc::new(
CircuitBreakerRegistry::new(5, 60).with_storage(storage.clone()),
);
cb_registry.load_from_storage().await;What decides whether the breaker helps or hurts in a shared engine is the key. My first version keyed it by handler name. Tenant A pointed http_request at a broken endpoint, the breaker tripped, and every tenant using http_request (which is nearly all of them) got rejected for a minute. The reliability feature had turned one customer outage into a platform outage.
#[derive(Clone, Debug)]
struct Key(TenantId, String);
pub struct CircuitBreakerRegistry {
breakers: DashMap<Key, CircuitBreakerState>,
default_threshold: u32,
default_cooldown_secs: u64,
storage: Option<Arc<dyn StorageBackend>>,
tracker: TaskTracker,
}The composite key creates a Rust problem. check runs on every step dispatch, and a naive DashMap<(TenantId, String), _> lookup means building an owned tuple, which means a String allocation, per check. The fix is the Borrow trick you use to look up a HashMap<String, _> with a &str, extended to a two-field key. A KeyRef(&TenantId, &str) struct and the owned Key both implement a small CircuitKey trait, Key implements Borrow<dyn CircuitKey>, and Hash and Eq are defined on the trait object so the borrowed and owned forms hash and compare identically. That last part is the requirement Borrow imposes and the one people miss. If the two forms disagree on hashing, lookups silently fail.
struct KeyRef<'a>(&'a TenantId, &'a str);
impl Hash for dyn CircuitKey + '_ {
fn hash<H: Hasher>(&self, state: &mut H) {
self.tenant_id().hash(state);
self.handler().hash(state);
}
}
pub fn check(&self, tenant_id: &TenantId, handler: &str) -> Result<(), u64> {
let search = KeyRef(tenant_id, handler);
let q: &dyn CircuitKey = &search;
if let Some(breaker) = self.breakers.get(q)
&& matches!(breaker.state, BreakerState::Closed | BreakerState::HalfOpen)
{
return Ok(()); // read lock only, no write, no allocation
}
// ... Open: check cooldown under get_mut; unknown key: insert default
}check is also careful about lock shape. A healthy breaker is answered under a DashMap read lock. Only an Open breaker takes the shard write lock, to move it to HalfOpen when the cooldown has elapsed. record_success returns early when the breaker is already Closed with zero failures, and record_failure returns early when it is already Open. Under load, nearly every call takes the read path.
Persisting Open state, and only Open
An in-memory breaker resets on restart. Picture a fleet restart in the middle of a dependency outage: every tripped breaker comes back Closed, and every node resumes hammering the thing that is down, at the moment it can least cope. So Open transitions are mirrored to a circuit_breakers table (migration 029) and reloaded at boot with their opened_at, which preserves the cooldown clock.
CREATE TABLE IF NOT EXISTS circuit_breakers (
tenant_id TEXT NOT NULL,
handler TEXT NOT NULL,
state TEXT NOT NULL,
failure_count INTEGER NOT NULL,
failure_threshold INTEGER NOT NULL,
cooldown_secs BIGINT NOT NULL,
opened_at TIMESTAMPTZ,
PRIMARY KEY (tenant_id, handler)
);
CREATE INDEX IF NOT EXISTS idx_circuit_breakers_open
ON circuit_breakers (state) WHERE state = 'open';Only Open rows get written. The reasons:
- Closed is the default. A row for every untripped breaker would mean writing a record to say nothing is wrong, for every (tenant, handler) pair that ever ran.
- HalfOpen is a probe owned by the live process. After a restart there is no process running that probe, so the state has no meaning. A persisted Open row whose cooldown has already elapsed is flipped to HalfOpen by the first
checkafter boot, which gives the same effect without storing it. - Open is the one state whose loss changes behaviour for the worse. When the breaker moves from Open to HalfOpen or Closed, the row is deleted so a crash cannot revive a breaker that had already recovered.
The write is off the hot path. A transition spawns the upsert or delete onto a tokio_util::task::TaskTracker, so check, record_failure, record_success, and reset stay synchronous. The tracker exists because the first version used a bare tokio::spawn, and a write in flight during runtime shutdown could be aborted halfway, leaving the table disagreeing with what the process had in memory. Now the server calls flush at shutdown, which closes the tracker and awaits every outstanding write.
Handlers that should not have a breaker
The breaker exists for external dependencies: HTTP APIs, LLM providers, plugin sidecars, gRPC services. Put one on a handler with no external dependency and you only get false positives. A fail step in a test suite would trip the breaker for fail. A send_signal that loses a race with the target completing would count as a failure. Neither should take down every instance using that handler for a minute.
pub fn is_breaker_tracked(handler: &str) -> bool {
!matches!(
handler,
"noop" | "log" | "sleep" | "fail" | "self_modify"
| "emit_event" | "send_signal" | "query_instance" | "human_review"
)
}I keep the skip-list here instead of as a property on each handler because it is policy that has to be reviewed as a set. When the list lives in one place, "which handlers bypass the breaker?" is a five-second answer. Scattered opt-outs mean grepping, and a wrong entry hides for a long time.
When a breaker is open the engine does not simply fail the step. breaker_preflight_at is a pure decision function with no storage writes. If the step declares a fallback_handler and that handler's own breaker is not open (or it is on the skip-list), the step definition is cloned with the handler swapped and dispatch continues. Otherwise the function returns a fire_at equal to now plus the remaining cooldown, and the caller moves the instance from Running back to Scheduled at that time. The decision is shared by both dispatch paths, the flat step loop in step_exec and the tree evaluator in step_block, so open-breaker behaviour cannot drift between them.
pub(crate) enum BreakerDecision<'a> {
Proceed(Cow<'a, StepDef>),
Defer { fire_at: DateTime<Utc> },
}
let Some(fb) = step_def.fallback_handler.as_deref() else {
let fire_at = now + Duration::seconds(remaining_secs as i64);
return BreakerDecision::Defer { fire_at };
};
// re-check the fallback's own breaker, then:
let mut cloned = step_def.clone();
cloned.handler = fb.to_string();
BreakerDecision::Proceed(Cow::Owned(cloned))Fair claiming with a per-tenant cap
Breakers handle a tenant whose dependency is failing. They do nothing about a tenant who is simply large. If the claim query takes rows in priority-then-time order, a tenant with 500,000 due instances fills every batch, and everyone else waits with no error anywhere. Just a queue that never reaches them.
The Postgres claim query caps rows per tenant using ROW_NUMBER() OVER (PARTITION BY tenant_id). Postgres will not allow FOR UPDATE SKIP LOCKED in the same query level as a window function, so the rows are locked in a CTE first and ranked in the outer query. The CTE over-selects limit * max_per_tenant rows, because if it locked only limit rows and one tenant owned all of them, the cap would trim the batch to almost nothing. The full story of that query, including the version that was wrong, is in the scheduler article linked below.
WITH locked AS (
SELECT id, tenant_id, priority, next_fire_at
FROM task_instances
WHERE (next_fire_at IS NULL OR next_fire_at <= $1) AND state = 'scheduled'
ORDER BY priority DESC, next_fire_at ASC NULLS FIRST
LIMIT $4 -- limit * max_per_tenant
FOR UPDATE SKIP LOCKED
), ranked AS (
SELECT id, tenant_id, priority, next_fire_at,
ROW_NUMBER() OVER (PARTITION BY tenant_id
ORDER BY priority DESC, next_fire_at ASC NULLS FIRST) AS rn
FROM locked
), winners AS (
SELECT id FROM ranked WHERE rn <= $3
ORDER BY priority DESC, next_fire_at ASC NULLS FIRST
LIMIT $2
)
SELECT task_instances.* FROM task_instances JOIN winners USING (id)One thing I should be clear about: the cap is opt-in. ORCH8_MAX_INSTANCES_PER_TENANT defaults to 0, which means no limit, and the query falls back to a plain LIMIT $2 FOR UPDATE SKIP LOCKED. A single-tenant deployment pays nothing for fairness it does not need. A multi-tenant deployment has to set it, and I would rather that be a visible configuration decision than a default that changes claim behaviour for everyone.
Tenant checks inside cross-workflow handlers
An engine that lets workflows spawn, signal, and query each other has created a way for one tenant to reach another. The enforcement has to be inside the handler, where the target instance is loaded, and not in a wrapper that a future code path can skip.
emit_eventspawns a child instance in the caller's tenant only. An optionaldedupe_keycan be scoped to the parent instance or to the whole tenant, and the dedupe row plus the child instance are written in one storage transaction viacreate_instance_with_dedupe, so a crash mid-operation cannot leave a dedupe row pointing at a child that was never created.send_signalloads the target, checkstarget.tenant_id == ctx.tenant_id, and only then callsenqueue_signal_if_active. That method re-reads the target state under a row lock (FOR UPDATEon Postgres, write-transaction semantics on SQLite) and refuses terminal targets in the same transaction as the insert. The first version checked state, then inserted, and a target could complete in between.query_instancereads same-tenant instances only. A missing target and a target in another tenant both return{ found: false }.
let target = match target {
Some(t) if t.tenant_id == ctx.tenant_id => t,
_ => return Ok(json!({ "found": false })),
};That last one is the kind of leak that gets through a review. The access check is correct and no data crosses the boundary, but if "does not exist" and "exists but is not yours" produce different responses, the response itself tells an attacker whether an instance ID is live in another tenant. send_signal does the same thing with its error: cross-tenant and not-found both return the message "target instance not found", and there is a test asserting the cross-tenant error does not mention the word terminal, because that would reveal the target exists and has finished.
Storage routing with no default backend
Logical isolation has a floor. Some tenants need their rows in a specific country, or on hardware their contract names, or simply not next to anyone else. That means routing a tenant to its own storage backend, and routing is the place where a convenient default turns into a data breach.
The TenantPartitionRouter in orch8-storage/src/tenant_partition.rs reads a placement record for every route and fails closed. A tenant with no record gets NotFound. A record naming a backend that this process did not register at boot gets Unsupported. Neither falls back to another partition. "We could not find your placement so we used the default one" is how a regulated tenant's data ends up in the wrong jurisdiction, and I would rather the request fail.
pub async fn route(&self, tenant_id: &TenantId) -> Result<RoutedTenantStorage, StorageError> {
let placement = self.placements.get_tenant_placement(tenant_id).await?
.ok_or_else(|| StorageError::NotFound {
entity: "tenant storage placement", id: tenant_id.to_string(),
})?;
let backend = self.backends.get(&placement.backend_id).ok_or_else(|| {
StorageError::Unsupported(format!(
"tenant '{}' is assigned to unregistered storage backend '{}'",
tenant_id, placement.backend_id))
})?;
Ok(RoutedTenantStorage { placement, backend: Arc::clone(backend) })
}Placement changes are epoch-fenced, the same way ownership is fenced everywhere else in the engine. The table (migration 078) has tenant_id as primary key, a backend_id, and a positive epoch. advance_tenant_placement is a conditional upsert: insert if absent, otherwise replace only when the proposed epoch is strictly greater than the stored one. A stale control-plane writer replaying an old placement gets a Conflict and changes nothing.
INSERT INTO tenant_storage_placements (tenant_id, backend_id, epoch, updated_at)
VALUES ($1, $2, $3, $4)
ON CONFLICT (tenant_id) DO UPDATE
SET backend_id = EXCLUDED.backend_id,
epoch = EXCLUDED.epoch,
updated_at = EXCLUDED.updated_at
WHERE EXCLUDED.epoch > tenant_storage_placements.epoch
-- rows_affected != 1 => StorageError::Conflictbackend_id is an operator-chosen name, validated to 1-128 characters of letters, digits, dots, dashes, and underscores. It is never a connection string. The placement table belongs in a highly available control-plane store and may be replicated more widely than credentials should travel. Connection details stay in each process's protected configuration, so the routing table is authoritative without being secret.
The router selects a backend. It does not copy data, coordinate dual writes, or make a multi-backend transaction atomic. Moving a tenant means quiescing writes, copying, validating, registering the destination on every serving node, and only then advancing the epoch. And as I said earlier, this component exists in the storage crate with its own tests and doc, and the server still boots against one backend. Wiring the router into the request path is the next piece of work, not finished work.
Designing tenant isolation into your own system
None of this is specific to a workflow engine. Any shared service has the same list of resources and the same question per resource.
- Write down the shared resources: connection pools, worker capacity, rate limits, caches, breakers, queues, scheduler batches, memory.
- For each one, decide whether the boundary is per tenant, global, or tiered. An undecided boundary is a global one.
- Put cross-tenant checks inside the operation that loads the foreign object. Middleware gets bypassed.
- Return the same response for "does not exist" and "not yours". Different errors are an existence oracle.
- Fail closed on routing. A default backend is convenient right up until the first compliance incident.
- Run the noisy-neighbour test on purpose. One tenant at ten times everyone else's load, and assert the others still hit their objectives. I did not do this early enough, and the handler-keyed breaker would have shown up on day one if I had.
The reason to do this early is that isolation is hard to retrofit. Every boundary has to exist at the point where the resource is consumed. With five call sites, adding a tenant dimension is an afternoon. With fifty, it is a migration.
Isolation Mechanisms and What Each Contains
| Mechanism | Boundary | Contains | Cost |
|---|---|---|---|
| Per-tenant circuit breaker | (tenant, handler) | One tenant's broken dependency | More breaker state in memory; a Borrow impl to avoid allocating per check |
| Persisted Open state | circuit_breakers table | Fleet restart resuming a stampede | Small crash window before the tracked write lands |
| Breaker skip-list | Handler class | False positives on internal steps | A list to keep correct |
| Fallback handler on open breaker | Step definition | A stalled step when an alternative exists | Per-step configuration |
| Per-tenant claim cap | Scheduler batch | One tenant monopolising throughput | Over-select in the claim CTE; opt-in via config |
| Tenant-scoped coordination | Inside emit_event, send_signal, query_instance | Cross-tenant spawns, signals, and reads | Checks on every coordination path |
| Uniform not-found response | Handler output | Cross-tenant existence oracle | Less specific error messages |
| Epoch-fenced placement | tenant_storage_placements | Stale routing after a move | Explicit migration procedure |
| No default backend | TenantPartitionRouter | Data landing in the wrong partition | Every tenant needs a record; not yet wired into the server |
None of these is expensive on its own. The expensive version is adding them to a system that assumed one shared pool for two years.
Related Matching Engine Guides
The claim query where per-tenant fairness is enforced, and the window-function restriction that broke the first version.
Why an unprotected failing dependency produces a queue of unresolved effect receipts.
The tenant boundary that every ownership transfer is validated against.
The execution model these isolation boundaries are built around.
Related reading
The pillar guide covers the full engine architecture: execution model, crate design, storage backends, and the operational decisions behind them.
Read the durable workflow engine architecture guide →