
An audit procedure for AWS architects and SREs that recounts the single-Region control planes, us-east-1 operation dependencies and SDK STS defaults AWS documents, from pages reviewed October 9, 2026. It covers IAM Identity Center during a Regional event, a CloudTrail replay script and data-plane replacements for common runbook steps.
At a glance
Key findings
- AWS documents 13 global services whose control plane runs in one Region, 10 in us-east-1 and 3 in us-west-2, while their data planes keep serving when those control planes fail. [1][2][3]
- Regional calls still reach us-east-1: 24 named S3 bucket-configuration operations, every CreateBucket and DeleteBucket, and resource creation in at least 13 services that writes Route 53 records. [1]
- Of 19 SDKs and tools in AWS's reference table, 9 list the global STS endpoint as their default and 6 fail requests with no Region set; the table lags Boto3 1.40.0 and AWS CLI v1 1.42.0, which moved to regional. [7][8][9]
- The STS global endpoint is served in-Region only for requests resolved by Amazon DNS in a Region enabled by default, so runbooks run from laptops or on-premises hosts still reach us-east-1. [10]
- An Identity Center additional Region serves sign-in, the access portal and session revocation, but not assignment changes or SCIM, so access must be provisioned before the event. [15][17]
Thirteen control planes that live in one Region
Thirteen AWS global services run their control plane in a single Region: ten in us-east-1 and three in us-west-2, according to AWS's Fault Isolation Boundaries whitepaper. A further set of ordinary Regional calls reaches us-east-1 indirectly. The whitepaper names 24 S3 bucket-configuration operations, plus every CreateBucket and DeleteBucket, and 13 services whose create, update or delete paths write Route 53 records through its control plane in us-east-1. A failover step that calls any of these can stall when that Region or that control plane is impaired, even while the recovery Region is healthy. [1]
The data planes of the same services are built to keep running through a control-plane failure. IAM keeps authenticating and authorizing existing principals in every Region, Route 53 keeps answering queries and running health checks, CloudFront keeps serving and failing over between origins, and Global Accelerator keeps routing on the dials and weights it already has. So the audit looks for one thing: steps that change configuration, or read it from a control plane, at the moment of recovery. [2][3]
It has three parts. Read the runbooks, pipelines and SDK configuration against the inventory below. Replay the last exercise from CloudTrail and list every write that landed in us-east-1 or us-west-2, plus every STS call served through the global endpoint. Then replace each finding with a data-plane operation, or write down that recovery waits for that Region. The diagram traces four common runbook steps, the dependency each takes today and its replacement.
Two cautions about the source. The whitepaper was first published on November 16, 2022 and its revision history ends with a minor update on February 9, 2023, so later features are missing from it; two of them, Route 53 accelerated recovery and ARC Region switch, change parts of the picture and are covered below. The counts here are this publication's own tallies of AWS's lists, made on October 9, 2026, and AWS calls its Route 53 list incomplete. [4][1]
Four runbook steps and where each one lands
Each step reaches a single-Region dependency today and has a data-plane replacement that does not. [1][10]

Source. Conceptual diagram based on the AWS Fault Isolation Boundaries whitepaper and the IAM User Guide, reviewed October 9, 2026. [1][2][3][10]
Method. Conceptual. Steps are a hypothetical runbook; each dependency and replacement is taken from the cited AWS guidance.
Accessible table and figure data
| Runbook step | Dependency today | Data-plane replacement |
|---|---|---|
| Get credentials | Global STS endpoint, served in us-east-1 from outside AWS | Regional STS endpoint |
| Move traffic | Route 53 record change, control plane in us-east-1 | ARC routing control on the recovery cluster |
| Add capacity | New load balancer records through Route 53 in us-east-1 | Capacity provisioned before the event |
| Grant access | IAM policy change, control plane in us-east-1 | Role created in advance with its policy |
| Runbook step | Dependency today | Data-plane replacement |
|---|---|---|
| Get credentials | Global STS endpoint, served in us-east-1 from outside AWS | Regional STS endpoint |
| Move traffic | Route 53 record change, control plane in us-east-1 | ARC routing control on the recovery cluster |
| Add capacity | New load balancer records through Route 53 in us-east-1 | Capacity provisioned before the event |
| Grant access | IAM policy change, control plane in us-east-1 | Role created in advance with its policy |
Thirteen control planes in two Regions
The whitepaper sorts global services into two groups. Partitional services exist once per partition with a control plane in one Region; some, such as IAM, also run an isolated data plane in every Region of the partition, while others, such as Network Manager, only orchestrate other services. Edge network services serve their data plane from points of presence that anyone on the internet can reach. In the aws partition that makes six partitional services and seven edge services. The right-hand column of the table is the one that matters for recovery design: whatever a runbook needs during the event has to come from it, or be done beforehand. [1][2][3]
Control plane here means more than writes. AWS counts every public IAM API as control plane, Access Advisor included but not Access Analyzer or IAM Roles Anywhere, and names ListAccounts among the Organizations control-plane operations. A reasonable reading is that a recovery script which discovers accounts with ListAccounts, or an application that looks up roles or groups through IAM at run time, depends on us-east-1 although it changes nothing. STS is the exception that matters most: AWS describes it as a data-plane-only service that does not depend on the IAM control plane. [2]
AWS's own summary of the October 2025 us-east-1 event shows a read carrying that dependency across Regions. From 11:47 PM on October 19 to 1:20 AM on October 20 (PDT), Redshift customers in every Region could not run queries with IAM user credentials, because a Redshift defect used an IAM API in us-east-1 to resolve user groups; customers who connected with local database users were unaffected. That dependency sat inside an AWS service, but the same shape is easy to build into application code. [5]
| Service | Control plane | Keeps working |
|---|---|---|
| AWS IAM | us-east-1 | Authentication and authorization in each Region |
| AWS Organizations | us-east-1 | SCPs and tag policies are still evaluated |
| AWS Account Management | us-east-1 | Account references in IAM policies |
| Route 53 Private DNS | us-east-1 | Resolution in each Region |
| Route 53 ARC | us-west-2 | Routing control state on the recovery cluster |
| AWS Network Manager | us-west-2 | Data planes of the networks it manages |
| Route 53 Public DNS | us-east-1 | DNS answers and health checks |
| Amazon CloudFront | us-east-1 | Caching, serving and origin failover |
| AWS WAF Classic for CloudFront | us-east-1 | Configured web ACLs and rules |
| AWS WAF for CloudFront | us-east-1 | Configured web ACLs and rules |
| ACM for CloudFront | us-east-1 | Existing certificates and automatic renewal |
| AWS Shield Advanced | us-east-1 | Configured protection and health checks |
| AWS Global Accelerator | us-west-2 | Routing, health checks, existing dials and weights |
Regional calls with a dependency on us-east-1
The whitepaper's third category is not a service but a set of operations you send to one Region that depend on another. For S3, 24 named bucket-configuration operations depend on us-east-1, among them PutBucketPolicy, PutBucketReplication, PutBucketEncryption, PutBucketVersioning and PutBucketPublicAccessBlock. Every CreateBucket and DeleteBucket call also depends on us-east-1 to keep bucket names unique, wherever the bucket lives. The control plane for S3 Multi-Region Access Points runs only in us-west-2 and leans on Global Accelerator there, Route 53 in us-east-1 and ACM; the access points' failover controls are separate, hosted in five Regions and usable as a data plane. [1]
Most names on that list are CloudTrail event names rather than API names, which helps the audit: CloudTrail records the PutBucketLifecycleConfiguration API as PutBucketLifecycle. Checked against S3's own list of bucket-level CloudTrail events, 23 of the 24 match. DeleteBucketLogging is not on it, so a filter on that name never fires; logging changes arrive as PutBucketLogging, which is on both lists. [6]
Route 53 is the dependency most often hidden in infrastructure code. A service that gives a new resource its own DNS name writes the records or health checks through the Route 53 control plane. AWS lists 13 such services: API Gateway REST and HTTP APIs, Elastic Load Balancing, PrivateLink VPC endpoints, Lambda function URLs, ElastiCache, OpenSearch Service, CloudFront, MemoryDB, Neptune, DAX, Global Accelerator, ECS with DNS-based service discovery through Cloud Map, and the EKS Kubernetes control plane. Creating an edge-optimized API Gateway endpoint also needs the CloudFront control plane, while EC2 hostnames in VPC DNS need neither. [1]
It follows that a CloudFormation stack or Terraform apply run in the recovery Region during failover is not Regional just because its provider points there. If the template creates a load balancer, an OpenSearch domain or an interface endpoint, part of its success depends on us-east-1. AWS's list of common anti-patterns covers the same ground:
- Editing a Route 53 record's value, or a weighted set's weights, to fail over.
- Creating or updating IAM roles and policies during failover, which AWS notes is usually unintended and may come from an untested failover plan.
- Changing Global Accelerator traffic dials by hand.
- Updating a CloudFront distribution's origin to move away from an impaired one.
- Provisioning recovery resources, such as load balancers and RDS instances, that need new Route 53 records. [1]
STS endpoints and what each SDK calls
Credentials are the first call in almost every runbook, and the whitepaper's statement about them is out of date: it says STS usage from the SDKs and CLI defaults to us-east-1. The AWS SDKs and Tools Reference Guide shows a split. Of the 19 SDKs and tools in its support table, 10 target a Regional STS endpoint by default and 9 the global endpoint sts.amazonaws.com. With no Region configured, 10 fall back to the global endpoint, 3 use the us-east-1 Regional endpoint (C++, JavaScript 3.x and .NET 4.x) and 6 fail the request: Go V2, Java 2.x, PHP, Ruby, Rust and Swift. The Java 2.x row adds that AssumeRole and AssumeRoleWithWebIdentity use the global endpoint when no Region is set. [1][7]
The table also disagrees with itself and with the SDK changelogs. Five rows (.NET 3.x, PHP, Boto3 and both PowerShell versions) give regional as the default setting but the global endpoint as the default target. The Boto3 changelog records that version 1.40.0 changed the default from legacy to regional, and AWS CLI v1 1.42.0 carries the same entry, yet the table still lists the CLI v1 default as legacy. This guide follows the changelogs for those two tools, because they record the code change. The version pinned in a runbook image therefore decides the behavior: an image built with an older Boto3 still calls the global endpoint by default. [7][8][9]
The global endpoint has also changed underneath these defaults. AWS now serves sts.amazonaws.com requests in the Region where they originate, but only in Regions enabled by default and only when the Amazon DNS server in a VPC resolved the name. A request from an opt-in Region, or one resolved by an ISP, a public resolver or any other DNS service, is still served in us-east-1. That is the path many runbooks take, from an operator laptop, an on-premises runner or a CI system outside AWS. CloudTrail records every global-endpoint call in us-east-1 whatever Region served it, and each such request carries aws:RequestedRegion set to us-east-1, which matters to any SCP that restricts Regions. [10]
For global-endpoint requests, STS adds endpointType and awsServingRegion under additionalEventData.RequestDetails, and the record's userAgent names the SDK or CLI that made the call. The fix is configuration for most tools: a Region plus sts_regional_endpoints = regional in every profile and runner image, and for SDKs without that setting, an explicit Region. Tools read different Region variables. The AWS CLI v2 reads AWS_REGION before AWS_DEFAULT_REGION, while the CLI v1 and Boto3 read only AWS_DEFAULT_REGION. [11][7][12]
# Shared config file (~/.aws/config) for a recovery runner
[profile dr-runbook]
region = us-west-2
sts_regional_endpoints = regional
# Environment for a container or pipeline step
AWS_REGION=us-west-2
AWS_DEFAULT_REGION=us-west-2
AWS_STS_REGIONAL_ENDPOINTS=regionalWhich STS endpoint AWS's table says each SDK uses
Nine of 19 SDKs and tools default to the global endpoint, and six fail when no Region is set. [7]

Source. AWS SDKs and Tools Reference Guide, table Support by AWS SDKs and tools on the AWS STS Regional endpoints page, reviewed October 9, 2026. Counts are rows in that table. [7]
Method. Each of the 19 rows was counted once per column: Default service client target STS Endpoint, and Service client fallback behavior. The table lags at least Boto3 1.40.0 and AWS CLI v1 1.42.0, which changed to regional defaults. [8][9]
Accessible table and figure data
| Behavior | Regional endpoint | Global endpoint | us-east-1 Regional | Request fails |
|---|---|---|---|---|
| Default STS endpoint | 10 | 9 | 0 | 0 |
| No Region configured | 0 | 10 | 3 | 6 |
| Behavior | Regional endpoint | Global endpoint | us-east-1 Regional | Request fails |
|---|---|---|---|---|
| Default STS endpoint | 10 | 9 | 0 | 0 |
| No Region configured | 0 | 10 | 3 | 6 |
Sign-in and Identity Center during a Regional event
AWS's post-event summary for October 19 and 20, 2025 shows the sign-in path failing in ways a Region map would not predict. From 11:51 PM to 1:25 AM PDT, console sign-in with IAM users failed more often, customers whose IAM Identity Center instance was in us-east-1 could not sign in through it, and root users and federation configured for signin.aws.amazon.com saw console sign-in errors in Regions outside us-east-1. STS in us-east-1 returned errors from 11:51 PM, recovered at 1:19 AM and degraded again from 8:31 to 9:59 AM after NLB health check failures. [5]
For SAML federation straight to IAM roles, the whitepaper's advice is to accept logins from several Regional sign-in endpoints, such as us-west-2.signin.aws.amazon.com, in every role trust policy, and to point the IdP at another Regional endpoint when the preferred one is impaired. The trust policy has to change in advance, because editing a role is itself an IAM control-plane call. AWS also recommends pre-provisioned break-glass users in case the IdP fails, and warns that an IdP hosted on AWS can be caught in the same event. [1][2]
IAM Identity Center added multi-Region replication on February 2, 2026, extended it to the Identity Center directory on July 28 and to opt-in Regions, AWS GovCloud (US) and the China Regions on September 30, 2026. An organization instance replicates workforce identities, permission sets, assignments and sessions from its primary Region to each additional Region, which gets its own access portal. Replication needs an external IdP or the Identity Center directory (not Active Directory) and a multi-Region customer managed KMS key in the same account. To get the full benefit, an external IdP must also accept multiple SAML assertion consumer service (ACS) URLs; AWS names Okta, Microsoft Entra ID, PingFederate, PingOne and JumpCloud as supporting them and Google Workspace as not. [13][14]
The table explains what a disruption leaves you. People can sign in, reach the portal, obtain role credentials and use the CLI, but only with assignments that existed before the primary failed: permission sets and assignments change only in the primary Region, and an additional Region cannot be promoted. SCIM does not run in additional Regions. It follows that deprovisioning someone in the IdP during the disruption stops new IdP sign-ins but does not reach Identity Center's copy of the user until the primary returns. Their existing sessions are closed by revocation, which administrators can perform in an additional Region and which replicates. [14][15][16][17]
Three details break failover in practice. CLI authorizations are not replicated, so each operator needs a profile and a login for the additional Region, made before the event. The additional Region has no custom alias and no awsapps.com start URL, only the instance URL under portal.amazonaws.com or the dual-stack app.aws domain. And AWS KMS does not synchronize key policies across a multi-Region key's Regions, so every policy change must be repeated on the replica. For an IdP with a single ACS URL, AWS's workaround is to switch that URL to the additional Region during the disruption, which needs a working IdP administrator. [17][16][18]
Replication does not replace break-glass access. AWS recommends both, because continuity through an additional Region still depends on the external IdP being healthy. [18]
| API | In an additional Region |
|---|---|
| Identity Center | Application management and instance reads; no permission set or assignment calls |
| Identity Store | Read operations only |
| OIDC | All operations |
| Access portal | All operations |
| SCIM | No operations |
# ~/.aws/config: one profile per Identity Center Region (AWS CLI v2)
[profile ops-primary]
sso_session = idc-primary
sso_account_id = 111122223333
sso_role_name = ReadOnly
region = us-west-2
[profile ops-additional]
sso_session = idc-additional
sso_account_id = 111122223333
sso_role_name = ReadOnly
region = us-west-2
[sso-session idc-primary]
sso_region = us-east-1
sso_start_url = https://ssoins-1111aaaa2222bbbb.portal.us-east-1.app.aws
sso_registration_scopes = sso:account:access
[sso-session idc-additional]
sso_region = eu-west-1
sso_start_url = https://ssoins-1111aaaa2222bbbb.portal.eu-west-1.app.aws
sso_registration_scopes = sso:account:access
# Authorize each session before an incident; authorizations do not replicate:
# aws sso login --sso-session idc-additionalRead the runbook, pipeline and SDK configuration
Start with everything the recovery path contains, not what an exercise happened to run: every runbook, the pipelines and Lambda functions it triggers, the infrastructure code it applies, and the images or hosts it runs on. Search them for calls that reach the inventory.
- CLI or SDK calls to IAM, Organizations, Account Management, Route 53, CloudFront, Global Accelerator, Shield and the ARC configuration API, reads such as
ListAccountsandDescribeClusterincluded. [2] - S3 calls from the bucket-configuration list,
CreateBucketandDeleteBucket, and WAF or ACM calls made for CloudFront, which go to us-east-1. - Templates or modules applied during failover that create load balancers, APIs, interface endpoints or anything else on the Route 53 list. [1]
- SDK versions,
regionandsts_regional_endpointsin every runner image, profile and function environment, checked against AWS's table and the SDK's changelog. - Identity paths: the start URL operators use, the SAML endpoint the IdP posts to, and whether a CLI profile exists for each Identity Center Region.
Replay the last exercise from CloudTrail
A static read misses calls that libraries and services make on your behalf, so compare it with what the last exercise actually sent. CloudTrail event history keeps 90 days of management events in each Region and needs no trail. LookupEvents accepts one lookup attribute per call and allows two requests per second per account and Region, which is enough for a few hours of exercise traffic. [19]
Since November 22, 2021, CloudTrail records IAM, CloudFront and global-endpoint STS events in us-east-1, and some other global services log in us-east-2 or us-west-2. Query us-east-1, us-east-2 and us-west-2 for every write made by the runbook roles, and each Region the runbook touched for bucket changes and resource creation. STS needs its own pass, because CloudTrail logs AssumeRole, AssumeRoleWithSAML and AssumeRoleWithWebIdentity as read-only events, so a write filter skips them. [20][11]
Stop and check three things before trusting the output. An exercise older than 90 days is outside event history, so query the trail's logs in S3 instead. A single-Region trail outside us-east-1 does not receive global service events. And an empty result shows only that this exercise made no such call; branches it skipped, such as a manual fallback, still depend on the static read. [19][20]
# Example: list control-plane writes and global-endpoint STS calls from an exercise.
# Read-only: needs cloudtrail:LookupEvents in each Region queried.
import json
from datetime import datetime, timezone
import boto3
from botocore.config import Config
START = datetime(2026, 9, 15, 14, 0, tzinfo=timezone.utc) # exercise window
END = datetime(2026, 9, 15, 18, 0, tzinfo=timezone.utc)
REGIONS = ["us-east-1", "us-east-2", "us-west-2", "eu-west-1"] # add every recovery Region
RUNBOOK = "arn:aws:sts::111122223333:assumed-role/dr-runbook/"
RETRY = Config(retries={"mode": "standard", "max_attempts": 10})
def lookup(region, key, value):
client = boto3.client("cloudtrail", region_name=region, config=RETRY)
pages = client.get_paginator("lookup_events").paginate(
LookupAttributes=[{"AttributeKey": key, "AttributeValue": value}],
StartTime=START, EndTime=END)
for page in pages:
for item in page["Events"]:
yield json.loads(item["CloudTrailEvent"])
for region in REGIONS:
for event in lookup(region, "ReadOnly", "false"):
arn = event.get("userIdentity", {}).get("arn", "")
if arn.startswith(RUNBOOK):
print(region, event["eventSource"], event["eventName"], arn)
for event in lookup("us-east-1", "EventSource", "sts.amazonaws.com"):
details = event.get("additionalEventData", {}).get("RequestDetails", {})
if details.get("endpointType") == "global":
print("global STS", details.get("awsServingRegion"), event["eventName"],
event.get("userAgent", ""))Replace each finding with a data-plane operation
Traffic movement has the most direct substitute. Route 53 health checks, and the routing changes that follow from them, keep working without the control plane, and ARC routing controls let an operator set those health checks by hand through the recovery cluster. Store the five Regional cluster endpoints instead of discovering them with DescribeCluster, and drive routing controls from the CLI or SDK rather than the console. ARC Region switch, which postdates the whitepaper, marks each of its API operations as data plane or not: GetPlanInRegion, ListPlansInRegion and StartPlanExecution are data-plane operations, while GetPlan and ListPlans are not, so a runbook calling the second pair takes a dependency it does not need. [3][2][21]
At the edge, prefer mechanisms that act on health rather than configuration. CloudFront origin groups fail over without a distribution update, within two limits that decide whether they fit: CloudFront still tries the primary origin first for every request, and fails over only GET, HEAD and OPTIONS requests. For writes or a full evacuation, a reasonable design keeps the origin's hostname behind health-checked Route 53 records, so the decision happens in the Route 53 data plane. Global Accelerator keeps routing on its existing health checks, dials and weights, so set those ahead of time and let health checks move the traffic. [22][3]
For IAM, create every role a recovery needs, attach its policies and test it before the event. Where an operator needs more privilege during an incident, the whitepaper suggests session tags from the IdP evaluated against aws:PrincipalTag, so neither the role nor the SCPs change. Data a recovery would otherwise fetch from a control plane, such as account lists, cluster endpoints and resource identifiers, belongs in a store read through its data plane, such as Parameter Store, DynamoDB or S3, optionally copied to a second Region. [2][1]
For S3 and provisioning, the replacement is to do the work earlier. Create recovery buckets with their policies, encryption, replication and versioning already set, and do not build a workflow that needs a particular bucket name to be free. Pre-provision load balancers, API Gateway endpoints and anything else on the Route 53 list. The October 2025 event shows why capacity belongs in the same category: EC2 instances launched before the event stayed healthy throughout, while new launches failed for hours. [1][5]
Recovery steps and their data-plane replacements
Every common control-plane step in a runbook has a replacement that is either a data-plane call or work done before the event. [1][2][3]

Source. Conceptual summary of the AWS Fault Isolation Boundaries whitepaper, ARC, CloudFront and IAM documentation, reviewed October 9, 2026. [1][2][3][21][22][10]
Method. Conceptual mapping. Each row pairs an operation AWS identifies as control plane or single-Region with the alternative AWS documents for it.
Accessible table and figure data
| Step | Reaches | Replace with |
|---|---|---|
| Edit a Route 53 record or weight | Route 53, us-east-1 | Health-check failover or ARC routing control |
| Discover ARC cluster endpoints | ARC configuration, us-west-2 | Stored list of five endpoints |
| Read a Region switch plan with GetPlan | Region switch control plane | GetPlanInRegion in the target Region |
| Create a load balancer, API or domain | Route 53, us-east-1 | Provision before the event |
| Create or edit an IAM role or policy | IAM, us-east-1 | Role made in advance; session tags |
| Edit an SCP or list accounts | Organizations, us-east-1 | Static SCPs; cached account list |
| Change a CloudFront origin | CloudFront, us-east-1 | Origin group failover |
| Change a Global Accelerator dial | Global Accelerator, us-west-2 | Endpoint health checks |
| Create a bucket or change its settings | S3 dependency on us-east-1 | Buckets configured in advance |
| Assume a role through sts.amazonaws.com | us-east-1 from outside AWS | Regional STS endpoint |
| Step | Reaches | Replace with |
|---|---|---|
| Edit a Route 53 record or weight | Route 53, us-east-1 | Health-check failover or ARC routing control |
| Discover ARC cluster endpoints | ARC configuration, us-west-2 | Stored list of five endpoints |
| Read a Region switch plan with GetPlan | Region switch control plane | GetPlanInRegion in the target Region |
| Create a load balancer, API or domain | Route 53, us-east-1 | Provision before the event |
| Create or edit an IAM role or policy | IAM, us-east-1 | Role made in advance; session tags |
| Edit an SCP or list accounts | Organizations, us-east-1 | Static SCPs; cached account list |
| Change a CloudFront origin | CloudFront, us-east-1 | Origin group failover |
| Change a Global Accelerator dial | Global Accelerator, us-west-2 | Endpoint health checks |
| Create a bucket or change its settings | S3 dependency on us-east-1 | Buckets configured in advance |
| Assume a role through sts.amazonaws.com | us-east-1 from outside AWS | Regional STS endpoint |
What still waits for us-east-1, and the order to fix the rest
Some work has no data-plane form. Creating hosted zones, accounts, IAM principals or CloudFront distributions, changing rules in a CloudFront web ACL and adding Shield Advanced protection all wait for their control plane. Anything of this kind left in the recovery path has to move before the event, or the plan has to say that recovery waits for that Region. WAF deserves its own decision, because an attack that calls for a rule change can arrive while us-east-1 is impaired, and existing rules are all that will be enforced. [2][3]
Route 53 accelerated recovery narrows one of these waits. For public hosted zones enabled in advance, Route 53 keeps a copy in us-west-2 and routes control-plane requests there within about 60 minutes after AWS detects that us-east-1 is impaired. During failover 13 API methods work, including ChangeResourceRecordSets, GetChange and ListResourceRecordSets; zones cannot be created or deleted, DNSSEC signing cannot change and private hosted zones are not covered. Changes accepted in us-east-1 but not yet copied are stranded and return NoSuchChange, so automation has to resubmit them, and enabling a zone can take several hours. SDK support arrived in Boto3 1.41.4. Its target is about an hour after AWS's own detection, which is a recovery objective for an API rather than static stability, so treat it as a backstop. [23][8]
With the inventory in hand, fix the findings in this order, cheapest and most common first:
- Set a Region and Regional STS in every runner, and raise pinned SDK versions past the change to a
regionaldefault. This removes the dependency an operator laptop takes on its first call. - Take IAM and Organizations calls, reads included, out of the recovery path: pre-create roles, use session tags for elevation and cache account lists.
- Replace DNS edits with health checks or ARC routing controls, then enable accelerated recovery on public zones as a backstop.
- Move resource creation, buckets and bucket configuration before the event, starting with anything on the Route 53 list.
- Replicate Identity Center to a distant Region, give every operator a second CLI profile, keep break-glass access and rehearse a sign-in through the additional Region.
Method and provenance
Source-led analysis of the AWS Fault Isolation Boundaries whitepaper, the AWS SDKs and Tools Reference Guide, IAM, IAM Identity Center, CloudTrail, S3, Route 53, CloudFront and ARC documentation, SDK changelogs and AWS's October 2025 post-event summary. Counts were made by the publication from AWS's lists. Sources were reviewed on October 9, 2026.
No AWS account, runbook or SDK was configured or exercised. Counts reflect AWS's lists and tables as published on October 9, 2026; AWS describes its Route 53 list as not exhaustive, and the SDK table lags some SDK releases.
AI assistance. AI assisted research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Global services, AWS Fault Isolation Boundaries whitepaper Amazon Web Services. Accessed .
- Appendix A: Partitional service guidance, AWS Fault Isolation Boundaries Amazon Web Services. Accessed .
- Appendix B: Edge network global service guidance, AWS Fault Isolation Boundaries Amazon Web Services. Accessed .
- Document revisions, AWS Fault Isolation Boundaries Amazon Web Services. Published . Accessed .
- Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region Amazon Web Services. Accessed .
- Amazon S3 CloudTrail events, Amazon S3 User Guide Amazon Web Services. Accessed .
- AWS STS Regional endpoints, AWS SDKs and Tools Reference Guide Amazon Web Services. Accessed .
- Boto3 changelog (versions 1.40.0 and 1.41.4) Amazon Web Services. Accessed .
- AWS CLI v1 changelog (version 1.42.0) Amazon Web Services. Accessed .
- AWS STS Regions and endpoints, IAM User Guide Amazon Web Services. Accessed .
- Logging IAM and AWS STS API calls with AWS CloudTrail, IAM User Guide Amazon Web Services. Accessed .
- AWS Region setting, AWS SDKs and Tools Reference Guide Amazon Web Services. Accessed .
- Document history, IAM Identity Center User Guide Amazon Web Services. Accessed .
- Using IAM Identity Center across multiple AWS Regions Amazon Web Services. Accessed .
- IAM Identity Center service APIs supported in additional AWS Regions Amazon Web Services. Accessed .
- Replicate IAM Identity Center to an additional Region Amazon Web Services. Accessed .
- Workforce access through an additional Region, IAM Identity Center Amazon Web Services. Accessed .
- Failover to an additional Region for AWS account access, IAM Identity Center Amazon Web Services. Accessed .
- LookupEvents, AWS CloudTrail API Reference Amazon Web Services. Accessed .
- CloudTrail concepts: global service events, AWS CloudTrail User Guide Amazon Web Services. Accessed .
- Region switch API operations, Amazon Application Recovery Controller Amazon Web Services. Accessed .
- Optimize high availability with CloudFront origin failover Amazon Web Services. Accessed .
- Enabling accelerated recovery for managing public DNS records, Amazon Route 53 Amazon Web Services. Accessed .