
A decision guide for architects choosing managed databases for multi-Region services, based on AWS, Google Cloud and Microsoft documentation and SLAs reviewed in October 2026. It gives a comparison matrix, a calculated conversion of published SLAs into monthly minutes, a decision tree and a test plan per replication model.
At a glance
Key findings
- DynamoDB MREC, Aurora Global Database failover and Cosmos DB session or weaker consistency replicate asynchronously; AWS puts MREC RPO at usually a few seconds and Microsoft puts Cosmos DB under 15 minutes. [1][5][14]
- Zero RPO needs a synchronous quorum: DynamoDB MRSC, Aurora DSQL multi-Region, Spanner dual-region or multi-region, or Cosmos DB strong consistency with one write Region. [1][8][12][14]
- Synchronous options cost features and geography: MRSC needs exactly three Regions and drops transactions, TTL and LSIs, Aurora DSQL cannot cross continents, and Cosmos DB blocks strong consistency beyond 5,000 miles by default. [1][7][14]
- Aurora Global Database switchover has zero RPO only between healthy clusters; unplanned failover uses
--allow-data-lossand leaves a snapshot of the old primary for reconciliation. [5] - Five-nines SLAs equal 0.432 minutes in a 30-day month against 4.32 for four nines, but they are credit commitments, measured differently, and none covers RPO. [3][10][13][17]
Consistency is a recovery decision
The consistency model of a multi-region database decides what happens to the last writes when a Region fails. With asynchronous replication, a write is acknowledged once it is durable in one Region, so a Region failure can strand writes that never left it: you lose them or reconcile them later. With synchronous quorum replication, a write is acknowledged only after another Region, or a majority of voting replicas, holds it, which is what makes a recovery point objective (RPO) of zero possible. Every managed option sits on one side of that line, and its documentation says which. [1][5][8][12][14]
The synchronous side is paid for on every write rather than during failures. Commit latency grows with the round trip to the farthest voting Region, the set of usable Regions shrinks, and some features disappear. The asynchronous side keeps local write latency and moves the cost to the failure, when someone has to decide which writes survived, which were lost and which conflicting versions win. [1][14]
Choose per dataset, starting from the recovery objective. A dataset that can lose its last few seconds of writes is cheaper and simpler to run on asynchronous replication. A dataset that cannot, such as an order ledger or a stock reservation, needs a synchronous system that fits its access pattern and latency budget. The sections below set out the documented behavior of DynamoDB global tables, Aurora Global Database, Aurora DSQL, Spanner and Azure Cosmos DB as reviewed on October 8, 2026, followed by a decision tree and a test plan.
What asynchronous replication loses
DynamoDB global tables default to multi-Region eventual consistency (MREC). Changes replicate asynchronously, typically within a second or less, and AWS states that the RPO equals the replication delay between replicas, usually a few seconds depending on the Regions. If the same item is written in two Regions, DynamoDB keeps the write with the latest internal timestamp for that item. A strongly consistent read returns the latest version only when the item was last written in the reader's own Region, and conditional writes evaluate against the local copy. Transactions are atomic only in the Region where they ran, so a reader elsewhere can briefly see part of a TransactWriteItems call. [1]
Aurora Global Database has one writer Region and up to 10 read-only secondary Regions, replicated at the storage layer with lag typically under a second. [4] It offers two ways to move the writer. A switchover, previously called managed planned failover, waits until the target secondary is fully synchronized, makes the old primary read-only and then promotes, so RPO is zero; AWS intends it for clusters that are healthy. A failover, requested with --allow-data-loss, does not wait, so writes not yet replicated are missing from the new primary. [5]
Aurora's failover procedure shows what recovery from asynchronous replication involves. Aurora tries to fence writes at the old primary's storage layer, but AWS calls this best effort and warns that writes might be momentarily accepted in the old Region, which is split brain. When the old Region returns, Aurora attempts a snapshot of its volume at the point of failure, named with the prefix rds:unplanned-global-failover-, so the missing writes can be recovered. That system snapshot follows the old cluster's backup retention period; copy it to a manual snapshot if reconciliation will take longer. [5]
Aurora PostgreSQL offers one way to bound the loss. The rds.global_db_rpo parameter, minimum 20 seconds, blocks commits on the primary whenever every secondary lags beyond the target. It trades write availability for a bounded loss window, and for global databases with only two Regions AWS recommends keeping the default in the secondary's parameter group, because otherwise transactions can pause after a failover while the old Region is rebuilt. [5]
Cosmos DB accounts in more than one Region publish an RPO of less than 15 minutes for session, consistent prefix and eventual consistency. Bounded staleness limits loss to K versions or T seconds, but for multi-region accounts the minimums are 100,000 write operations or 300 seconds, so it is a bounded option rather than a near-zero one. [14]
# Planned move between healthy Regions: waits for sync, RPO zero
aws rds switchover-global-cluster \
--region us-east-1 \
--global-cluster-identifier example-global \
--target-db-cluster-identifier arn:aws:rds:us-west-2:111122223333:cluster:example-secondary
# Unplanned recovery: does not wait, unreplicated writes can be lost
aws rds failover-global-cluster \
--region us-west-2 \
--global-cluster-identifier example-global \
--target-db-cluster-identifier arn:aws:rds:us-west-2:111122223333:cluster:example-secondary \
--allow-data-lossWhat synchronous quorums cost
DynamoDB multi-Region strong consistency (MRSC), generally available since June 30, 2025, replicates each change synchronously to at least one other Region before the write succeeds, and AWS states that it supports an RPO of zero. [1][2] Its constraints are specific. A table must span exactly three Regions, as three replicas or as two replicas plus a DynamoDB-managed witness that serves no reads or writes, chosen from 15 supported Regions, including combinations across continents. It can only be created from an empty table, replicas cannot be added later, and TTL, local secondary indexes and transactions are not supported. The consistency mode is fixed at creation. A write to an item that another Region is modifying fails with ReplicatedWriteConflictException and can be retried. [1]
Aurora DSQL, generally available since May 27, 2025, applies the same pattern to PostgreSQL-compatible SQL. A multi-Region cluster has two peered Regional endpoints that both accept reads and writes, plus a witness Region that stores only the encrypted transaction log and has no endpoint. AWS describes replication as always synchronous, with no failover operation and no loss from replication lag. [8][11] Two constraints shape the choice. Peered clusters must stay within one Region set (North America, Asia Pacific or Europe), because cross-continent clusters are not supported. Concurrency control is optimistic: a transaction that loses a conflict fails at commit with SQLSTATE 40001 and code OC000, so the application has to retry whole transactions. [7][9]
Spanner expresses the tradeoff through instance configurations, and both multi-Region shapes require the Enterprise Plus edition. A dual-region configuration keeps two read-write replicas and one witness in each of two Regions in one country, and commits need at least two replicas in each Region while it runs in dual-region mode. If one Region fails, writes can fail until the quorum is switched to the surviving Region. A manual failover usually completes within one minute; Google-managed failover can take up to 45 minutes. Google states that dual-region provides zero RPO. A multi-region configuration has five voting replicas across two read-write Regions and a witness Region; each write needs a leader-Region replica plus any two others. Voting Regions are generally under a thousand miles apart, and Google describes the cost as a small increase in write latency. [12]
Cosmos DB strong consistency commits each write in every Region of the account, or in a majority when dynamic quorum drops unresponsive Regions from accounts with three or more. Microsoft gives multi-region strong write latency as twice the round trip between the two farthest Regions plus 10 milliseconds at the 99th percentile. It blocks strong consistency by default when Regions span more than 5,000 miles (8,000 kilometers) and does not allow it with multiple write Regions. A two-Region strong account loses write availability when either Region fails unless you take the failed Region offline. [14][15]
Managed options compared
The matrix puts the documented commit path, Region-failure behavior and main constraints side by side. Two rows are easy to misread. Aurora Global Database's zero-RPO switchover is a planned operation for healthy clusters; in an unplanned outage the operation is failover, with a nonzero RPO. Cosmos DB with multiple write Regions keeps accepting writes elsewhere when one Region fails, but recent writes made in the failed Region can be unavailable until it recovers, when they are merged by the container's conflict resolution policy. [5][15]
Recovery time varies as much as recovery point. Aurora says the promoted secondary typically takes the primary role within a few minutes, and that switchover on qualifying engine versions typically restores reads and writes in under 30 seconds. [5] Cosmos DB service-managed failover can take one hour or more, which is why Microsoft recommends a forced failover, taking the Region offline, when write availability is needed quickly. [15] Spanner's Google-managed dual-region failover can take up to 45 minutes. [12] DynamoDB MRSC and Aurora DSQL have no promotion step; clients move to a healthy Regional endpoint. [1][8]
Where each option commits a write
Only the synchronous configurations give zero RPO in an unplanned Region failure; Aurora's zero-RPO switchover is for healthy clusters. [1][5][8][12][14][15]

Source. AWS, Google Cloud and Microsoft documentation reviewed October 8, 2026. [1][4][5][7][8][12][14][15][16]
Method. Documented behavior summarized per configuration; cells paraphrase the cited pages. Excludes pricing, backups and self-managed replication.
Accessible table and figure data
| Option | Commit and RPO | When one Region fails | Constraints to check |
|---|---|---|---|
| DynamoDB MREC | Async; RPO usually a few seconds | Shift traffic; last writer wins | Transactions atomic only per Region |
| DynamoDB MRSC | Sync to a second Region; RPO zero | Other replicas keep serving | Exactly 3 Regions; no transactions, TTL or LSIs |
| Aurora Global Database | Async; lag typically under 1 second | Failover can lose unreplicated writes | Zero RPO only in planned switchover |
| Aurora DSQL multi-Region | Sync; no replication lag | Peer endpoint keeps reads and writes | One Region set; retry OCC conflicts |
| Spanner dual-region | Quorum in both Regions; RPO zero | Writes can fail until quorum fails over | Enterprise Plus; two Regions in one country |
| Spanner multi-region | Leader replica plus 2 of 4 voters | Quorum can form from two or three Regions | Enterprise Plus; writes favor leader Region |
| Cosmos DB strong, one write Region | Global majority; RPO zero | Two-Region accounts lose write availability | Blocked beyond 5,000 miles by default |
| Cosmos DB session or weaker | Async; RPO under 15 minutes | Failover can lose unreplicated writes | Managed failover can take an hour or more |
| Cosmos DB multiple write Regions | Async; strong consistency not allowed | Other Regions keep writing | Merge by last writer wins or custom procedure |
| Option | Commit and RPO | When one Region fails | Constraints to check |
|---|---|---|---|
| DynamoDB MREC | Async; RPO usually a few seconds | Shift traffic; last writer wins | Transactions atomic only per Region |
| DynamoDB MRSC | Sync to a second Region; RPO zero | Other replicas keep serving | Exactly 3 Regions; no transactions, TTL or LSIs |
| Aurora Global Database | Async; lag typically under 1 second | Failover can lose unreplicated writes | Zero RPO only in planned switchover |
| Aurora DSQL multi-Region | Sync; no replication lag | Peer endpoint keeps reads and writes | One Region set; retry OCC conflicts |
| Spanner dual-region | Quorum in both Regions; RPO zero | Writes can fail until quorum fails over | Enterprise Plus; two Regions in one country |
| Spanner multi-region | Leader replica plus 2 of 4 voters | Quorum can form from two or three Regions | Enterprise Plus; writes favor leader Region |
| Cosmos DB strong, one write Region | Global majority; RPO zero | Two-Region accounts lose write availability | Blocked beyond 5,000 miles by default |
| Cosmos DB session or weaker | Async; RPO under 15 minutes | Failover can lose unreplicated writes | Managed failover can take an hour or more |
| Cosmos DB multiple write Regions | Async; strong consistency not allowed | Other Regions keep writing | Merge by last writer wins or custom procedure |
What the published SLAs allow
Each provider commits to more availability for its multi-Region configuration than for one Region: 99.999% for DynamoDB global tables, Aurora DSQL multi-Region clusters, Spanner dual-region and multi-region instances, and Cosmos DB multi-region reads and multiple write Regions, against 99.99% for the single-Region equivalents. [3][10][13][17] Converted at (1 minus SLA) times 43,200 minutes in a 30-day month, 99.99% allows 4.32 minutes and 99.999% allows 0.432 minutes, about 26 seconds.
An SLA is a service-credit commitment. The provider credits part of the bill if measured uptime falls below the threshold; the percentage is not a prediction of availability. Measurement also differs. DynamoDB averages availability across five-minute intervals by server error rate, Spanner counts only five-minute periods with more than a five percent error rate, and Cosmos DB subtracts an average hourly error rate. [3][13][17] The minutes in the chart are an equivalent full outage for comparison, not a budget to plan against.
Three details matter when quoting these numbers. Aurora's SLA covers Multi-AZ clusters at 99.99% and has no separate commitment for Global Database, so a global database does not carry a five-nines SLA. [6] The DynamoDB global tables SLA applies only if every table in the Region is part of a global table for the whole billing cycle and you make reasonable attempts to fail over, and Aurora DSQL's multi-Region SLA likewise expects you to use the other peered cluster. [3][10] Spanner regional instances in Mexico and Stockholm carry 99.95%. [13] No SLA covers RPO: the DynamoDB SLA page does not distinguish MREC from MRSC, so two tables with the same five-nines commitment can lose very different amounts of data in the same failure. [3]
Five nines allows about 26 seconds a month
Multi-Region configurations commit to 99.999%, equal to 0.432 minutes in a 30-day month, against 4.32 minutes at 99.99%. [3][6][10][13][17]

Source. Calculated from the AWS DynamoDB, Aurora and Aurora DSQL SLAs, the Cloud Spanner SLA and the Microsoft Online Services SLA of October 1, 2026, reviewed October 8, 2026. [3][6][10][13][17]
Method. Minutes = (1 minus SLA percentage / 100) x 43,200. SLAs are service-credit commitments, not availability predictions, and each provider measures uptime differently; minutes are a full-outage equivalent. Aurora Global Database has no separate SLA; Spanner regional is 99.95% in Mexico and Stockholm.
Accessible table and figure data
| Service and configuration | SLA (%) | Minutes per 30-day month |
|---|---|---|
| DynamoDB standard table (99.99%) | 99.99 | 4.32 |
| DynamoDB global tables (99.999%) | 99.999 | 0.432 |
| Aurora Multi-AZ cluster (99.99%) | 99.99 | 4.32 |
| Aurora DSQL single-Region (99.99%) | 99.99 | 4.32 |
| Aurora DSQL multi-Region (99.999%) | 99.999 | 0.432 |
| Spanner regional (99.99%) | 99.99 | 4.32 |
| Spanner dual-region (99.999%) | 99.999 | 0.432 |
| Spanner multi-region (99.999%) | 99.999 | 0.432 |
| Cosmos DB single region (99.99%) | 99.99 | 4.32 |
| Cosmos DB single region with zones (99.995%) | 99.995 | 2.16 |
| Cosmos DB multi-region reads (99.999%) | 99.999 | 0.432 |
| Cosmos DB multiple write regions (99.999%) | 99.999 | 0.432 |
| Service and configuration | SLA (%) | Minutes per 30-day month |
|---|---|---|
| DynamoDB standard table (99.99%) | 99.99 | 4.32 |
| DynamoDB global tables (99.999%) | 99.999 | 0.432 |
| Aurora Multi-AZ cluster (99.99%) | 99.99 | 4.32 |
| Aurora DSQL single-Region (99.99%) | 99.99 | 4.32 |
| Aurora DSQL multi-Region (99.999%) | 99.999 | 0.432 |
| Spanner regional (99.99%) | 99.99 | 4.32 |
| Spanner dual-region (99.999%) | 99.999 | 0.432 |
| Spanner multi-region (99.999%) | 99.999 | 0.432 |
| Cosmos DB single region (99.99%) | 99.99 | 4.32 |
| Cosmos DB single region with zones (99.995%) | 99.995 | 2.16 |
| Cosmos DB multi-region reads (99.999%) | 99.999 | 0.432 |
| Cosmos DB multiple write regions (99.999%) | 99.999 | 0.432 |
Conflict and reconciliation
Wherever two Regions accept writes asynchronously, the database picks a winner and the application inherits the result. DynamoDB MREC and Cosmos DB both default to last writer wins. DynamoDB applies it per item by internal timestamp. Cosmos DB uses the system _ts property by default and, in the API for NoSQL, accepts a custom numeric conflict resolution path; a delete always beats a concurrent insert or replace. [1][16] Last writer wins suits data where the newest value is the correct value, such as a user preference. It is wrong for counters, balances and reservations, where two concurrent decrements must both apply.
The Cosmos DB API for NoSQL also offers a custom policy: a merge stored procedure that runs exactly once per conflict, with failures written to the conflicts feed for the application to resolve. The policy can only be set when the container is created. [16] The conflict feed matters after a forced failover of a single-write-Region account too: when the old write Region comes back, writes it never replicated can be read from the feed and written back by application logic. [15]
The dependable patterns live in the application. Give each item a home Region and route its writes there, so cross-Region conflicts cannot arise for that item. DynamoDB's documentation suggests a Region-of-origin attribute for filtering stream records, and the same kind of attribute can record each item's home Region. [1] Attach a client-generated request identifier to each write so a retry after ReplicatedWriteConflictException or OC000 cannot apply twice. [1][9] Keep data whose invariants cannot survive last writer wins in a synchronous system and let the rest replicate asynchronously.
Reconciliation after an asynchronous failover needs its own runbook. For Aurora it starts from the rds:unplanned-global-failover- snapshot, which AWS provides so missing data can be recovered. [5] Restore it into a separate cluster, find transactions committed after the last replicated point and replay those still valid against the new primary. In a hypothetical order service, that means comparing order identifiers between the restored snapshot and the new primary and resubmitting only orders that are missing and have not since been superseded. The restore itself needs the same discipline as any point in time recovery: an isolated copy and acceptance checks before data returns to production.
Choose per dataset
Most services hold data with different tolerances. Take a hypothetical checkout service with a session cart, a product catalog and an order ledger: only the ledger may need zero RPO. Putting all three in a synchronous store charges cross-Region commit latency to data that does not need it, while putting the ledger in an asynchronous store turns every regional failure into a reconciliation job.
The tree asks about RPO first, because no tuning turns asynchronous replication into zero loss. It then asks whether more than one Region writes the same item, because that forces a merge rule, and whether writes need multi-row SQL transactions, which MRSC does not support and Aurora DSQL and Spanner do. The last two questions check Region geography and measured write latency, the constraints most likely to rule out a synchronous option late in a project. [1][7][12][14] Answer these for each dataset:
- The largest write loss the business accepts for this dataset, in seconds or transactions, and who accepted it.
- Whether any item can be written from two Regions at once, and what the correct merged value would be.
- The Regions where users and compute run, checked against the 15 MRSC Regions, the Aurora DSQL Region sets, the Spanner configurations or the Cosmos DB 5,000-mile limit for strong consistency.
- The 99th percentile write latency budget, measured from each application Region to the farthest voting Region.
Choose the replication model from the recovery point
Decide RPO first, then concurrent writers, transactions, geography and latency.

Source. Conceptual decision aid based on AWS, Google Cloud and Microsoft documentation. [1][5][7][12][14]
Method. Conceptual ordering of documented constraints. It does not cover cost, self-managed databases or data residency rules.
Accessible table and figure data
| Question | Yes, then | No, then |
|---|---|---|
| Can this dataset lose its last seconds of writes when a Region fails? | asynchronous replication is acceptable; ask about concurrent writers | require a synchronous quorum; ask about transactions |
| Can two Regions update the same item at the same time? | define the merge rule, then use MREC or Cosmos DB multiple write Regions | use one write Region: Aurora Global Database or Cosmos DB single write Region |
| Do writes need multi-row SQL transactions? | consider Aurora DSQL or Spanner dual-region or multi-region | DynamoDB MRSC and Cosmos DB strong are also candidates |
| Do your Regions fit the provider's supported Region sets and distance limits? | ask about measured write latency | pick another provider or keep this data in one write Region |
| Is 99th percentile write latency to the farthest voting Region within budget? | adopt it and drill Region isolation | keep only the data that needs zero RPO synchronous |
| Question | Yes, then | No, then |
|---|---|---|
| Can this dataset lose its last seconds of writes when a Region fails? | asynchronous replication is acceptable; ask about concurrent writers | require a synchronous quorum; ask about transactions |
| Can two Regions update the same item at the same time? | define the merge rule, then use MREC or Cosmos DB multiple write Regions | use one write Region: Aurora Global Database or Cosmos DB single write Region |
| Do writes need multi-row SQL transactions? | consider Aurora DSQL or Spanner dual-region or multi-region | DynamoDB MRSC and Cosmos DB strong are also candidates |
| Do your Regions fit the provider's supported Region sets and distance limits? | ask about measured write latency | pick another provider or keep this data in one write Region |
| Is 99th percentile write latency to the farthest voting Region within budget? | adopt it and drill Region isolation | keep only the data that needs zero RPO synchronous |
Test the failure you chose
Test the failure the model implies, not a generic Region outage. For an asynchronous store, the question is whether you can see the loss window and reconcile it. Alarm on ReplicationLatency for MREC tables and AuroraGlobalDBRPOLag for Aurora secondaries, and use Aurora switchover as the routine drill, since failover with --allow-data-loss against production is a real data-loss event. [1][5]
For a synchronous store, the question is whether the application keeps working while one Region is isolated and whether write latency stays inside budget in that state. AWS Fault Injection Service can pause replication to and from a DynamoDB global table replica, for MREC and MRSC alike, and can raise connection error rates in one Region of an Aurora DSQL cluster. [1][8] For Cosmos DB drills, Microsoft suggests disabling service-managed failover, running a forced failover and then restoring the setting. [15] Spanner dual-region publishes a quorum health timeline metric for the decision of when to fail over. [12]
Observe the clients as well as the database. Aurora recommends the global writer endpoint and a client DNS cache TTL as low as 5 seconds, because write fencing is best effort and a client still resolving the old writer is how split brain happens. [5] Record when each client actually stopped writing to the old Region, not only when the database reported promotion.
Two rules and a sequence
Two rules settle most of the arguments. If you cannot state the merge rule for an item, do not let two Regions write it. If you cannot accept the loss window for a dataset, do not replicate it asynchronously, whatever its SLA percentage says. Beyond those rules, work from the data outward in this order:
- Classify each dataset by tolerated write loss and record the owner who accepted that figure.
- Mark items that more than one Region can write and define their merge rule before choosing a database.
- Place zero-RPO data in a synchronous configuration that fits its access pattern and Regions, then measure 99th percentile write latency from every application Region.
- Place the rest on asynchronous replication, alarm on replication lag and write the reconciliation runbook, including snapshot retention.
- Quote SLA percentages as credit terms only and set recovery objectives from your own drills.
- Repeat the Region isolation drill whenever the topology, Regions or consistency settings change.
Method and provenance
Source-led technical analysis of AWS, Google Cloud and Microsoft product documentation and service level agreements, with an original comparison matrix, a calculated SLA conversion and a conceptual decision tree. Sources were reviewed on October 7 and 8, 2026.
No cloud account, database or fault injection experiment was used. Behavior, limits, Region lists and SLA terms are bounded to the cited pages as of the review date and can change; SLA conversions express credit thresholds, not measured availability.
AI assistance. AI assisted research synthesis, drafting, calculation checks and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- How DynamoDB global tables work Amazon Web Services. Accessed .
- Amazon DynamoDB global tables with multi-Region strong consistency is now generally available Amazon Web Services. Published . Accessed .
- Amazon DynamoDB Service Level Agreement Amazon Web Services. Published . Accessed .
- Using Amazon Aurora Global Database Amazon Web Services. Accessed .
- Using switchover or failover in Amazon Aurora Global Database Amazon Web Services. Accessed .
- Amazon Aurora Service Level Agreement Amazon Web Services. Published . Accessed .
- What is Amazon Aurora DSQL? Amazon Web Services. Accessed .
- Resilience in Amazon Aurora DSQL Amazon Web Services. Accessed .
- Concurrency control in Aurora DSQL Amazon Web Services. Accessed .
- Amazon Aurora DSQL Service Level Agreement Amazon Web Services. Published . Accessed .
- Amazon Aurora DSQL is now generally available Amazon Web Services. Published . Accessed .
- Regional, dual-region, and multi-region configurations Google Cloud. Accessed .
- Cloud Spanner Service Level Agreement (SLA) Google Cloud. Published . Accessed .
- Consistency level choices in Azure Cosmos DB Microsoft. Accessed .
- Reliability in Azure Cosmos DB Microsoft. Accessed .
- Conflict resolution types and resolution policies in Azure Cosmos DB Microsoft. Accessed .
- Service Level Agreement for Microsoft Online Services, October 1, 2026 Microsoft. Published . Accessed .