Staged Rollout Strategies for Multi-Tenant Schema Migrations
In the previous sections, we learned how to define target groups and deploy migrations to multiple tenant databases. In this section, we will explore deployment rollout strategies - a powerful feature that gives you fine-grained control over how migrations are applied across your tenant databases.
Overview
When deploying schema migrations to multiple tenant databases, you often need more control than simply applying migrations to all targets at once. Common requirements include:
- Canary deployments - Validate changes on a small subset of tenants before a broader rollout
- Regional rollouts - Deploy to US-West first, then US-East, then EU, minimizing blast radius
- Priority ordering - Migrate enterprise customers before SMB, or process tenants alphabetically
- Parallel execution - Speed up deployments by running migrations concurrently within each stage
- Error resilience - Log failures and continue with remaining tenants instead of stopping entirely
Atlas's deployment block addresses all these needs by organizing targets into groups with configurable
execution order, parallelism, and error handling.
The deployment Block
The deployment block defines a rollout strategy that can be referenced by one or more environments.
Basic Syntax
deployment "<name>" {
// Variables passed from the env block
variable "<var_name>" {
type = <type> // string, bool, number, etc.
default = <value> // Optional default value
}
// Groups define execution stages
group "<group_name>" {
match = <expr> // Boolean expression to filter targets
order_by = <expr> // Expression to sort targets within group
parallel = <number> // Max concurrent executions (default: 1)
on_error = FAIL | CONTINUE // Error handling mode
depends_on = [group.<other_group>] // Groups that must complete first
}
}
Connecting to an Environment
To use a deployment strategy, reference it in your env block using the rollout block:
env "prod" {
for_each = toset(var.tenants)
url = urlsetpath(var.url, each.value)
rollout {
deployment = deployment.staged
vars = {
name = each.value
}
}
}
Group Matching Behavior
By default, groups are evaluated in the order they appear in the configuration file. When a target matches multiple groups,
the first matching group wins - the target is assigned to it and skipped by subsequent groups. You can override
the execution order using the depends_on attribute.
This allows you to define specific groups first (e.g., canary, internal) followed by a catch-all group for remaining targets:
deployment "staged" {
variable "name" {
type = string
}
// First: Internal tenants (matched first by position)
group "internal" {
match = startswith(var.name, "internal-")
}
// Second: Canary tenants
group "canary" {
match = startswith(var.name, "canary-")
parallel = 10
depends_on = [group.internal]
}
// Last: Catch-all for remaining targets (no match = all unmatched)
group "rest" {
depends_on = [group.canary]
}
}
Group Attributes
match
A boolean expression that determines which targets belong to this group. Targets matching multiple groups are assigned to the first matching group by file position.
group "internal" {
match = startswith(var.name, "my-company-") || var.name == "internal-test"
}
order_by
Controls the execution order of targets within a group. Targets are sorted by this expression in ascending order. You can use a single expression or an array for multi-level sorting.
group "alphabetical" {
match = var.tier == "FREE"
order_by = var.name // Execute tenants alphabetically
}
group "by_region_then_name" {
order_by = [var.region, var.name] // Sort by region first, then by name
}
parallel
Maximum number of concurrent migrations within the group. Default is 1 (sequential execution).
group "free_tier" {
match = var.tier == "FREE"
parallel = 10 // Run up to 10 migrations concurrently
}
on_error
Defines behavior when a migration fails within the group:
FAIL(default) - Stop the group's execution immediately on first errorCONTINUE- Log the error and proceed with remaining targets in the group
group "non_critical" {
match = var.tier == "FREE"
on_error = CONTINUE // Log failures but continue with other tenants
}
Deployment Duration and Parallelism
Deployment time grows with the number of targets. Every target in a group is a full apply cycle
against its own database: Atlas connects to it, takes an
advisory lock, reads the applied versions from its
revisions table, and executes the files that are
still pending. When the environment enables the pre-apply drift check,
each target also fetches the expected state for its latest revision from the Atlas Registry, is
inspected, and is diffed against that state before any migration file runs. This work is fast: the
drift check in the sample transcript takes 12.4ms. But it runs for every target, so on a large fleet
it adds up and sets the minimum deployment time.
Targets already at the latest version
A target that is already at the latest version is a no-op: Atlas tracks applied versions and reports
No migration files to execute without executing any statement. Atlas Cloud records such a run as a
NO_ACTION event, and
atlas migrate status
reports Already at latest version for that target. In the
declarative flow, the equivalent output is
Schema is synced, no changes to be made.
This is what makes a rollout re-runnable and a converged fleet cheap to deploy to: synced targets do nothing beyond the check that they are up to date, and the run only executes statements on the targets that are behind.
Sizing parallel
parallel caps how many targets a group processes concurrently and defaults to 1, so a group of
100 tenants with parallel = 10 runs at most ten migrations at a time. Size that number against the
capacity of the instance the group's targets live on, not against how many tenants the group holds.
In a database-per-tenant architecture, compute and other resources are shared across tenants, so ten
concurrent migrations are ten concurrent workloads on an instance that is also serving the tenants
that are not being migrated.
When the fleet spans several instances, fetch the instance each tenant lives on as
tenant metadata, pass it in the
rollout vars, and match on it, the same way the tiered example below matches on var.tier. Each
group then covers one instance and its parallel value is sized for that instance alone.
Atlas takes an advisory lock while applying migrations to prevent conflicting parallel executions. If
two Atlas processes target the same database, one acquires the lock and runs while the other waits
for it to be released, failing if the lock is not freed within the --lock-timeout window (10s by
default). Targets that are separate schemas of the same database contend for that lock, since the
lock name defaults to atlas_migrate_execute, so raising parallel alone does not make their
migrations run concurrently.
Scope the lock per target with the --lock-name option (or migration.lock_name in the atlas.hcl
file), an Atlas Pro feature available after running atlas login. When coordination
is guaranteed externally, for example by a deployment system that runs a single migration job, the
--skip-lock flag skips acquiring the lock. See
Controlling Advisory Locks.
Recovering from a Partial Rollout
A rollout can end with the fleet in a mixed state. A group that hits an error under the default
on_error = FAIL stops there, and a group with on_error = CONTINUE logs the failure and proceeds,
so failures can be spread across the group. Recovery is the same in both cases: find the targets that
failed, fix what made them fail, then re-run the fleet.
Finding the failed targets
Once deployments are reported to the Atlas Registry, the information the Atlas Cloud UI shows is also available from the CLI, which is what you want when triaging from a script, a CI job, or an AI agent. List the failed deployment events:
atlas cloud migration list --status FAILED
ID REPO TYPE ENV DATABASE TARGETS VERSION STATUS
105 inventory SCHEMA production (multiple) 2/3 20260515090000 FAILED
106 payments MIGRATION_DIRECTORY eu-west prod-eu - 20260512141500 FAILED
--------------------------------
Page: 1 Page Size: 6 Total: 18
A fleet deployment shows (multiple) under DATABASE and a succeeded/total ratio under TARGETS,
so 2/3 means two databases succeeded before the failure. Drill into one event by ID:
atlas cloud migration describe --id 105
To look at the fleet by target instead of by run, list the databases with their sync status
(SYNCED, PENDING, or FAILED), optionally filtered by environment:
atlas cloud database list --env-name production
atlas cloud database describe --id 1 prints a single target with its current version and last
deployment time. The full output of each command is shown in
Inspecting Failures from the CLI
and Inspecting Deployments from the CLI.
Re-running the fleet
Fix the cause first, usually by correcting the problematic data in the target database, then re-run the same command that started the rollout:
atlas migrate apply --env prod
There is no need to compute the list of tenants that still need the change. As described above,
targets that already reached the latest version are a no-op and report
No migration files to execute, so the run only executes statements on the targets that are behind.
A canary group that is already synced is a no-op on the retry as well. The declarative flow behaves
the same way, see
Partial rollouts and re-runs.
Confirming the fleet is green
The rollout is done when every target is on the latest version. From the Registry side, list the
databases again and check that no row is left in PENDING or FAILED. From the apply side, a re-run
that prints No migration files to execute for every target reports the same thing. For a single
target, check it directly with
atlas migrate status,
which prints Already at latest version when nothing is pending.
Practical Examples
Canary Deployment Pattern
Deploy to a single canary tenant first, then roll out to everyone else:
deployment "canary" {
variable "name" {
type = string
}
group "canary" {
match = var.name == "canary-tenant"
}
group "rest" {
parallel = 5
depends_on = [group.canary]
}
}
Tiered Rollout by Customer Plan
Roll out to internal tenants first, then free tier (with high parallelism), then paid customers (more carefully):
data "sql" "tenants" {
url = var.management_url
query = <<SQL
SELECT name, tier
FROM tenants
SQL
}
env "prod" {
for_each = toset(data.sql.tenants.values)
url = urlsetpath(var.url, each.value.name)
migration {
dir = "atlas://my-app"
}
rollout {
deployment = deployment.tiered
vars = {
name = each.value.name
tier = each.value.tier
}
}
}
deployment "tiered" {
variable "name" {
type = string
}
variable "tier" {
type = string
}
// Internal tenants first - one at a time
group "internal" {
match = startswith(var.name, "my-company-")
}
// Free tier next - parallelize aggressively.
group "free" {
match = var.tier == "FREE"
parallel = 10
on_error = CONTINUE
depends_on = [group.internal]
}
// Paid customers last - more conservative
group "paid" {
parallel = 3
on_error = FAIL
depends_on = [group.free]
}
}