Skip to content
Cloud Security DeskSearch
Menu

Technical guideResilience

Plan for an outage at your CDN or edge provider

When the CDN, WAF and edge provider fails, traffic can move to a regional WAF, a second CDN or the origin, and each keeps less protection. Decide early whether a second path is worth securing.

Published
Sources checked
Next review
Reading time
16 minutes
Coverage
Microsoft Azure · Cloudflare · Amazon Web Services · Google Cloud
A broad blue canopy with a scalloped lower edge stretches across the top of the frame over a row of six small white origin blocks on a ground line, while a single narrow amber path curves beneath the canopy from the left edge to the right.
Conceptual illustration: the edge provider covers every origin at once, and any second path runs narrower, beneath it.

A decision framework for architects fronting public applications with Azure Front Door, CloudFront, Cloudflare or Cloud Armor, based on Microsoft's 2026 Front Door high-availability guide, Traffic Manager documentation and the primary 2025 incident reports reviewed October 9, 2026. It maps which edge functions survive on each failover path, calculates DNS switching time from documented settings, and gives a decision tree, origin trust tests and the conditions for staying down.

At a glance

Key findings

  • Microsoft's own architecture guidance says most workloads will not need a second global ingress path, and warns that a carelessly built one can make availability worse. [1]
  • Microsoft's January 2026 Front Door high-availability guide uses a manual break-glass switch with Traffic Manager health checks off, because Traffic Manager probes originate only from US Azure regions and cannot judge the global health of an anycast edge. [2]
  • The bypass path runs a different WAF policy: Front Door and Application Gateway both inspect the first 128 KB of a request body by default, only Application Gateway can be raised to 2 MB, a new Application Gateway policy starts in detection mode, and a second CDN applies its own rules. [2][16]
  • The 2025 edge outages took more than page delivery: Front Door's October 29 failure caused DNS errors and impaired the Azure portal and support case creation, and Cloudflare's November 18 failure stopped Turnstile and broke most Access sign-ins. [4][5]
  • At the 300-second Traffic Manager TTL in Microsoft's guide, a resolver can keep the old answer for five minutes after the switch, three times the calculated 100-second detection at default probe settings. [2][3]

Decide before the edge fails

When a global CDN, WAF and edge platform fails, traffic has three places to go: a regional WAF you operate, such as Azure Application Gateway; a second CDN; or straight to the origin. Each keeps less of what the edge did. A regional WAF keeps managed rules from the same product family but gives up caching, anycast distribution and edge DDoS absorption. A second CDN keeps a cache and a global network but enforces another vendor's rules. The direct path keeps nothing, and on Cloudflare it also publishes the origin address the edge was hiding. [1][2][10]

Whether to build any of them starts where Microsoft's architecture guidance starts: in most situations you won't need a second ingress path, and one built carelessly can lower availability by adding components, control planes and configuration that must stay in step. [1] Microsoft's January 28, 2026 Front Door high-availability guide, written for customers who build one anyway, describes a manual break-glass design: Traffic Manager in front, health checks off, and an operator who switches an endpoint when outside monitoring shows the edge has failed. [2]

The security cost is easy to underrate because the second path is never dormant. Disabling a Traffic Manager endpoint removes it from DNS answers but leaves the service behind it running. [3] The gateway, second CDN property or origin listener stays on the internet every day, which is why Microsoft warns that an attacker who finds an unprotected secondary path can use it while the primary still has its WAF. [1] Build a second path only if it will be secured and tested like the first; otherwise plan to stay down on purpose, and never improvise a direct-to-origin bypass mid-incident.

What an edge outage takes with it

Three primary incident reports from late 2025 show what customers lost; their delivery mechanics are analyzed elsewhere. Azure Front Door's October 29, 2025 outage ran from 15:41 UTC to 00:05 UTC the next day. Customers saw connection timeouts and DNS resolution errors, because Front Door's internal DNS service runs on the same edge sites that were crashing. The affected list included the Azure portal, parts of Microsoft Entra ID, Defender External Attack Surface Management, Sentinel threat intelligence and customers' ability to open support cases. [4]

Microsoft's own services show failover under pressure. The Entra and Intune portals and Azure AD B2C failed over. The Azure portal moved away from Front Door at 17:26 UTC, 1 hour 45 minutes after impact began, and parts of it with no fallback of their own, such as Marketplace, kept failing. [4] Failover is per dependency: a page served through the standby path still breaks if one of its components stays on the failed one.

The edge also stopped accepting changes. After mitigation Microsoft blocked all Front Door customer configuration changes at the Azure Resource Manager level until November 5, 2025, and propagation stayed slower for every operation, WAF changes and cache purges included. [4] A customer who needed an emergency WAF rule that week had to put it somewhere other than the edge.

Cloudflare's November 18, 2025 failure stopped core traffic from 11:20 UTC until about 14:30. Turnstile failed to load, so most users could not sign in to the Cloudflare dashboard, whose login page uses it. Access authentication failed for most users until Cloudflare applied a bypass at 13:05, though existing sessions continued. Sites on the older FL proxy kept serving, but every request got a bot score of zero, so rules that block bots turned away large numbers of real visitors. [5]

On December 5, 2025 the failure lasted about 25 minutes, from 08:47 to 09:12 UTC, and touched about 28 percent of Cloudflare's HTTP traffic. The customers hit were those on the FL1 proxy with the Cloudflare Managed Ruleset deployed, whose every request returned HTTP 500. The trigger came from protective work: while raising the WAF body buffer from 128 KB to 1 MB for CVE-2025-55182 in React Server Components, Cloudflare switched off an internal WAF testing tool through its global, non-gradual configuration system, and that change hit a bug in FL1. [6]

Read for planning, the reports say three things. The management plane, sign-in and support channel can fail with the data plane. A security feature can be the failure condition, or keep running on wrong inputs. And a switch that waits on a human decision and a DNS cache helps only with the long outages.

Map what the edge does for you

Build the inventory from the edge configuration. Microsoft's question list works for any provider: whether you rely on caching and whether the origin can carry the load without it; whether the rules engine routes or rewrites requests; whether the WAF protects the application; whether you restrict traffic by IP address or geography; who issues the TLS certificates; how the origin is restricted to edge traffic and whether it accepts traffic from anywhere else; and whether clients depend on the edge's HTTP/2 support. [1]

Add the functions the 2025 reports exposed: edge-hosted authentication such as Cloudflare Access, challenge widgets such as Turnstile that your sign-in pages call, bot scores referenced in your rules, and the provider's dashboard and API as the place emergency changes would be made. [5] Give each line an owner and a note of what replaces it on every candidate path, or a statement that nothing does.

The matrix records documented answers and marks reasoned ones. A regional WAF keeps rule continuity and gives up caching and edge absorption; a second CDN keeps caching and absorption and gives up rule continuity; the direct path gives up all of it. Edge-hosted sign-in stays with the original vendor on every path, which is easy to miss when failover is treated as a routing change.

Figure 01

What each failover path keeps of the edge

A regional WAF keeps rule continuity, a second CDN keeps caching and absorption, and the direct path keeps nothing; edge-hosted sign-in stays with the original vendor on all three. [1][2][9][10][11][12]

Matrix of nine edge functions against three failover paths: regional WAF, second CDN and direct to origin. Managed WAF rules, caching, DDoS absorption, certificates, rewrites, sign-in, origin trust, body inspection and Private Link origins each show what survives on each path.

Source. Conceptual comparison based on Microsoft's Front Door high-availability and global routing guidance, Azure WAF limits, and Cloudflare's origin, proxy status, Access and Turnstile documentation, reviewed October 9 and 10, 2026. [1][2][16][9][10][11][12]

Method. Conceptual matrix. Regional WAF cells follow Microsoft's Application Gateway scenario; second CDN cells follow its alternative CDN scenario; direct to origin follows Cloudflare's DNS-only behavior. The cold cache and the origin certificate requirement are reasoned inferences, not documented statements.

Accessible table and figure data
Figure 1 accessible table
Edge functionRegional WAFSecond CDNDirect to origin
Managed WAF rulesSame family, separate policy and exclusionsAnother vendor's rules, tuned separatelyNone
Body inspection depth128 KB default, up to 2 MB on Application GatewayThe vendor's own limitNone
CachingNone; the origin takes every requestYes, starting coldNone
Edge DDoS absorptionPlatform baseline unless DDoS Protection addedThe vendor's networkOrigin network baseline only
TLS certificateSame BYO certificate, Key Vault or uploadSame BYO certificate, uploadedOrigin must serve the public name
Edge rules and rewritesRebuilt as rewrite rule setsRebuilt in the vendor's rulesMoved into the application
Edge-hosted sign-in and challengesStill on the edge vendorStill on the edge vendorStill on the edge vendor
Origin trustNew source range and credentialNew secret, token or mTLSOrigin open to the internet
Private Link originsNeeds a private endpoint in the VNetCannot reach themNot applicable
Figure 1 accessible table
Edge functionRegional WAFSecond CDNDirect to origin
Managed WAF rulesSame family, separate policy and exclusionsAnother vendor's rules, tuned separatelyNone
Body inspection depth128 KB default, up to 2 MB on Application GatewayThe vendor's own limitNone
CachingNone; the origin takes every requestYes, starting coldNone
Edge DDoS absorptionPlatform baseline unless DDoS Protection addedThe vendor's networkOrigin network baseline only
TLS certificateSame BYO certificate, Key Vault or uploadSame BYO certificate, uploadedOrigin must serve the public name
Edge rules and rewritesRebuilt as rewrite rule setsRebuilt in the vendor's rulesMoved into the application
Edge-hosted sign-in and challengesStill on the edge vendorStill on the edge vendorStill on the edge vendor
Origin trustNew source range and credentialNew secret, token or mTLSOrigin open to the internet
Private Link originsNeeds a private endpoint in the VNetCannot reach themNot applicable

Three places traffic can go

In Microsoft's first scenario the target is a regional WAF. A primary Traffic Manager profile with weighted routing and health checks off points at Front Door and, disabled until needed, at a nested secondary profile. That profile uses performance routing across Application Gateway WAF_v2 instances in the regions with meaningful user volume, probes them over HTTPS and answers with a 300-second TTL, standing in for Front Door's latency-based routing. Application Gateway does not cache, runs only where you deploy it, and costs money while idle: about $200 to $400 a month for WAF_v2 at minimal capacity, by Microsoft's estimate. [2]

The second scenario uses a second CDN: one weighted profile with Front Door enabled and the other CDN's edge hostname disabled. The standby needs the same custom domain, origin configuration, caching rules, compression and query string handling, the same bring-your-own certificate, and WAF rules matched as closely as its product allows. [2] It keeps a global network and a cache, although a cache that carried no traffic before the switch can be expected to start cold and send the origin a burst of misses.

Direct to origin is the improvised third option, which Microsoft's guide does not offer. On Cloudflare it usually means switching a proxied record to DNS only, after which Cloudflare answers with the origin's real address and stops applying WAF rules, caching and its other products to that hostname. [10]

Private origins narrow the choice. Microsoft does not recommend Front Door Private Link origins for either scenario, because other CDNs cannot reach them and Application Gateway needs its own virtual network and private endpoint configuration to do so. [2] A Cloudflare Tunnel origin has no publicly routable address at all. [9] It follows that the most private origins are the hardest to give a second path, and adding one can mean making a private origin public.

Security controls the bypass path drops

The WAF changes even when the brand does not. Microsoft's guide pairs Default Rule Set 2.1 on Front Door with CRS 3.2 or 4.0 on Application Gateway, but the Application Gateway rule set reference lists no CRS 4.0, and DRS 2.2 was announced as generally available on both on March 17, 2026. [2][17][18] Even on one version they are separate policies, so match the Application Gateway rule set to Front Door's, test it independently and document every exclusion on both. [2] Both inspect the first 128 KB of a request body by default. Application Gateway with CRS 3.2 or DRS can be raised to 2 MB and in prevention mode blocks bodies over its size limit; Microsoft does not document what Front Door does with the rest of a larger body. [16][2] Moving from CloudFront to an Application Load Balancer takes AWS WAF from a default 16 KB of body, configurable to 64 KB, to a fixed 8 KB. [14] The same payload can be blocked on one path and pass on the other. A new Application Gateway WAF policy also starts in detection mode, so a standby never switched to prevention logs attacks instead of blocking them. [2]

Volumetric protection changes too. An origin that accepts only Front Door traffic is shielded by Front Door from layer 3 and layer 4 DDoS attacks. [1] An Application Gateway public IP has the Azure platform's baseline protection, which Microsoft says lacks workload-specific tuning, telemetry, cost protection and availability guarantees; its network security groups filter packets without stopping volumetric attacks, and Microsoft suggests Azure DDoS Protection on those addresses for production. [2]

Rules keyed to the client address need a second look on every path, because behind any proxy the WAF sees the proxy's address unless told to read a header. Google Cloud Armor, for example, exposes origin.user_ip, filled from a header named in userIpRequestHeaders such as True-Client-IP, and falling back to origin.ip when that header is missing or invalid. [13] Through a second CDN, allowlists, geographic rules and rate limits may see only the CDN's addresses. On a direct path, a header-derived client address is whatever the client sends, unless the network layer limits who can connect.

Edge-hosted sign-in does not move with the traffic. Turnstile requires your server to call Cloudflare's Siteverify endpoint for every token [12], so a sign-in form that uses it depends on Cloudflare on every path, and its owner must decide in advance whether a failed verification blocks the sign-in or skips the challenge. An application behind Cloudflare Access should validate the Cf-Access-Jwt-Assertion header against the account's signing key. [11] If it does, a direct path fails closed. If it trusts that Access ran at the edge and checks nothing, a direct path publishes an unauthenticated application, which is reason enough never to improvise a bypass for anything behind edge authentication.

Origin trust that works on both paths

Each edge has its own proof that a request came through it, and none accepts another provider's proof. Front Door origins should combine the AzureFrontDoor.Backend service tag with a check of the X-Azure-FDID header against the profile's identifier, because other Azure customers' Front Door profiles use the same addresses. [7] CloudFront origins pair a secret origin custom header with the AWS-managed prefix list in the load balancer's security group; AWS says to treat the header like a credential and rotate it by adding a second header, updating the listener rule, then removing the first. [8] With Cloudflare's Authenticated Origin Pulls, Cloudflare's own certificate proves only that a request came from Cloudflare's network, so stricter checks need your own uploaded certificate. [9]

For origins that must accept both paths, Microsoft lists CDN-agnostic controls: token-based origin authentication with HMAC or signed URLs, mutual TLS, custom origin headers and IP address filtering. [2] The editorial recommendation is to give each path its own credential, such as a client certificate from your own CA for each edge or a separate secret header per path, to allow each path's source ranges at the network layer, and to refuse any request whose credential and source range do not match. Separate credentials let you revoke one path after a leak or a vendor compromise without breaking the other.

Then test the failure Microsoft names: an origin accidentally opened to other paths, including other customers' Front Door profiles. [1] Each of these requests should behave as stated:

  • Sent straight to the origin address with no header or client certificate: refused.
  • Sent through a different profile, account or zone on the same edge provider: refused.
  • Carrying the previous, rotated secret: refused.
  • Sent through the standby path with its own credential while that path is disabled in DNS: served.

DNS, certificates and the switch

The switching record has to live where the failed edge cannot reach it. In Microsoft's design the custom hostname is a CNAME to Traffic Manager, which returns either the Front Door endpoint or the standby. [2] If the switching record is hosted by the edge provider, the switch depends on a dashboard and control plane that the November 18 report shows can be out of reach. [5] Traffic Manager cannot sit at the zone apex through a CNAME, so an apex name needs alias records from a provider such as Azure DNS. [2] Traffic Manager is then a dependency of its own, and Microsoft advises planning for a prolonged Traffic Manager outage by pointing DNS straight back at Front Door. [1]

Timing comes mostly from caches. Microsoft recommends TTLs of 300 to 600 seconds, sets 300 on both Traffic Manager profiles, and says to lower the hostname's CNAME TTL to 60 to 300 seconds at least 24 hours before moving it onto Traffic Manager, warning that propagation typically takes 5 to 10 minutes and can take 48 hours. [2] Cloudflare fixes proxied records at a 300-second TTL [10], so a DNS-only bypass waits on resolvers too. For the health-based secondary profile, Traffic Manager's defaults of a 30-second interval, three tolerated failures and a 10-second timeout mark an endpoint Degraded about 100 seconds after the first failed probe; fast probing at 10 seconds with a 9-second timeout cuts that to 39. [3] The cached answer, not detection, sets the floor.

Microsoft's guide is explicit that the primary profile should not fail over automatically: Traffic Manager probes originate only from US-based Azure regions, so against an anycast edge they almost always reach US points of presence and leave the rest unverified. [2] Traffic Manager also answers as if every endpoint were online when all of them are degraded, so a probe failure that hits both paths moves no traffic. [3] The global guide adds that mission-critical solutions require automated failover where possible, while telling global workloads to turn endpoint monitoring off. [1] The two fit if the trigger is automated but the probe is not Traffic Manager's: outside-in synthetic checks from the regions where users are, feeding a pre-approved script that disables the Front Door endpoint and that an operator can also run by hand. [1][2] That is the reading followed here.

Certificates have to be portable before the outage. Front Door managed certificates cannot be exported, so Microsoft requires bring-your-own certificates on both paths, with a PFX and private key for Application Gateway, held in Key Vault or uploaded directly. [2] Managed validation also assumes the hostname's CNAME points directly at Front Door, which stops being true once Traffic Manager is in the chain. [1] CloudFront accepts certificates only when requested or imported in the us-east-1 Region of AWS Certificate Manager [8], so a shared certificate is installed and renewed in two places, and Microsoft's guide makes testing renewal on every platform part of the design. [2]

Figure 02

After the switch, the cached answer sets the wait

At Microsoft's 300-second TTL, cached answers can outlast default Traffic Manager detection threefold; a manual switch adds the decision time, which no setting controls. [2][3]

Stacked horizontal bar chart in seconds. Health-based failover between gateway regions: 39 seconds of fast-probe detection plus a 300-second TTL, total 339; 100 seconds of default detection plus 300, total 400. Break-glass switch with no automated detection: 60-second TTL example, 300-second Traffic Manager TTL, and 600 seconds at the top of the recommended range.

Source. Calculated from Microsoft's Azure Front Door high-availability guide (ms.date January 28, 2026) and Traffic Manager endpoint monitoring documentation, both reviewed October 9, 2026. [2][3]

Method. Probe detection = tolerated failures x probing interval + probe timeout, the time from the first failed probe to the fourth consecutive failure in Microsoft's walkthrough: 3 x 10 + 9 = 39 and 3 x 30 + 10 = 100. Break-glass rows have no automated detection; the time to notice and decide is excluded because no source documents it. TTL is the longest a resolver may honor a cached answer; resolvers that keep answers longer, and the wait for the first failed probe, are excluded. [2][3]

Accessible table and figure data
Figure 2 accessible table
Switching method and settingAutomated probe detection (s)Cached answer TTL (s)
Health-based between gateway regions, 10 s fast probing39300
Health-based between gateway regions, 30 s default probing100300
Break-glass switch, 60 s TTL example060
Break-glass switch, 300 s Traffic Manager TTL0300
Break-glass switch, 600 s top of recommended range0600
Figure 2 accessible table
Switching method and settingAutomated probe detection (s)Cached answer TTL (s)
Health-based between gateway regions, 10 s fast probing39300
Health-based between gateway regions, 30 s default probing100300
Break-glass switch, 60 s TTL example060
Break-glass switch, 300 s Traffic Manager TTL0300
Break-glass switch, 600 s top of recommended range0600

Decide whether a second path is worth building

Microsoft prices a second path in four parts: running the standby, the operational complexity of every feature that exists on one path and not the other, extra CNAME lookups, and engineering time taken from other work. [1] Set that against the 2025 durations of about 25 minutes, about three hours and more than eight hours. [4][5][6] A second path covers only what remains of an outage after detection, decision and the TTL. In a hypothetical case with 15 minutes to detect and decide and a 300-second TTL, users reach the standby for at most the last five minutes of a 25-minute incident, but for more than seven hours of an eight-hour one.

Microsoft also says feature parity with Front Door should not be a strict requirement, and that the second path should carry what business continuity needs, even in limited capacity. [1] That makes partial failover the realistic shape: static content and read-only pages on the standby, sign-in and state-changing operations either fully protected there or deliberately left on the primary. The obvious choice is wrong in at least three common cases:

  • A high-value authenticated API looks like the first candidate for failover, but if the standby WAF has not been tested against it, failing over trades an availability incident for an exposure.
  • A static marketing site looks too unimportant to protect, yet it is the cheapest case for a second CDN: cacheable, no sign-in and a small rule set to keep in step.
  • An application behind edge authentication gains nothing from failover unless sign-in has another path, and becomes dangerous on a direct path if it does not validate the edge's token.
Figure 03

Decide whether to build a second ingress path

Every question that ends in no points to fixing a prerequisite or staying down, not to a weaker second path. [1][2]

Decision tree with six questions: whether an outage of several hours costs more than a year of standby and rehearsal, whether the origin can serve uncached peak traffic, whether a second WAF can be tested every quarter, whether the origin can authenticate both paths without Private Link, whether sign-in or bot rules depend on the edge vendor, and whether an operator can switch DNS from a control plane the edge does not host.

Source. Conceptual decision aid based on Microsoft's global routing redundancy guidance and Front Door high-availability guide, reviewed October 9, 2026. [1][2]

Method. Conceptual ordering of documented tradeoffs, cheapest disqualifier first. Answer each question for the specific application or hostname, not for the whole estate. [1][2]

Accessible table and figure data
Figure 3 accessible table
QuestionYes, thenNo, then
Would an outage of several hours cost more than a year of standby and rehearsal?continue to the capacity questionwrite a stay-down runbook instead
Can the origin serve peak traffic with no cache in front of it?a regional WAF path is an optiononly a second CDN fits, and its cache starts cold
Can a WAF on the second path be tuned and tested every quarter?continue to origin trustdo not build it; an untested path is a standing bypass
Can the origin authenticate both paths without Private Link?continue to sign-in dependenciesredesign origin access before adding a path
Do sign-in, challenges or bot rules depend on the edge vendor?give them a fallback or accept that sign-in stays downcontinue to the switch
Can a named operator switch DNS from a control plane the edge does not host?build the path and rehearse it every quarterfix DNS hosting and access first
Figure 3 accessible table
QuestionYes, thenNo, then
Would an outage of several hours cost more than a year of standby and rehearsal?continue to the capacity questionwrite a stay-down runbook instead
Can the origin serve peak traffic with no cache in front of it?a regional WAF path is an optiononly a second CDN fits, and its cache starts cold
Can a WAF on the second path be tuned and tested every quarter?continue to origin trustdo not build it; an untested path is a standing bypass
Can the origin authenticate both paths without Private Link?continue to sign-in dependenciesredesign origin access before adding a path
Do sign-in, challenges or bot rules depend on the edge vendor?give them a fallback or accept that sign-in stays downcontinue to the switch
Can a named operator switch DNS from a control plane the edge does not host?build the path and rehearse it every quarterfix DNS hosting and access first

Rehearse the switch

Microsoft's guide asks for failover and failback to be tested in non-production first and then every quarter. Before moving production DNS, it suggests editing the hosts file on a non-production workstation so the custom domain resolves to the standby, and it asks for WAF rule sets to be audited for consistency with every exclusion documented. [2] The global guide's validation questions make good acceptance criteria: traffic moves when the primary is unavailable, both paths carry production load with little warning, and the degraded state exposes nothing new. [1]

A security rehearsal does more than load the home page. Run the same WAF test set through both paths and compare decisions rule by rule. Repeat the origin refusal tests. Confirm that sign-in works on the standby, or fails the way its owner chose. Check that the standby WAF's logs reach the same detection pipeline. Time the switch from the decision to the first request seen on the standby, from resolvers in more than one region.

Run the switch from the CLI or REST API, not the portal: the October 29 outage impaired the Azure portal, and Microsoft's advice for that case is to manage resources programmatically. [4][1] The permission to switch is also the permission to send every user elsewhere, because Microsoft.Network/trafficManagerProfiles/externalEndpoints/write lets its holder add an external endpoint or update any property of an existing one, not only its status. [15] Grant it through a narrow role to a small named group, test the role outside production and alert on every use. After failback, rotate anything the incident exposed, including an origin address published by a DNS-only change. [9]

Example custom role definition for az role definition create, scoped to one placeholder resource group. Warning: externalEndpoints/write can also add an endpoint or repoint an existing one, so assigning this role hands over control of all traffic to the domain. Verify the switch commands in a test subscription before relying on it.
{
  "Name": "Edge failover operator (example)",
  "IsCustom": true,
  "Description": "Example: enable or disable Traffic Manager external endpoints during an edge outage. Cannot create, delete or reconfigure profiles.",
  "Actions": [
    "Microsoft.Network/trafficManagerProfiles/read",
    "Microsoft.Network/trafficManagerProfiles/externalEndpoints/read",
    "Microsoft.Network/trafficManagerProfiles/externalEndpoints/write"
  ],
  "NotActions": [],
  "AssignableScopes": [
    "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-edge-failover"
  ]
}

When to stay down instead

Staying down is a decision, and often the safer one. Choose it when the only path left is direct to origin and the origin relies on the edge for authentication, WAF or DDoS protection; when the standby WAF has not been tested against the current application this quarter; when the origin cannot carry uncached peak traffic, so failover moves the outage instead of ending it; or when the remaining outage is likely to be shorter than the decision plus the TTL. Publishing an origin address is not easily undone: Cloudflare advises rotating origin IPs after onboarding because historical DNS records are kept [9], and an emergency DNS-only change leaves the same trail.

A stay-down plan still needs preparation: provider health alerts routed to the right people, programmatic access to every console the team would use, a communication channel that does not run through the failed provider, and a list of edge changes to apply once the provider accepts changes again. After its October 29 outage, Front Door did not accept customer configuration changes until November 5. [4]

Method and provenance

Source-led decision analysis of Microsoft's Front Door high-availability and global routing guidance, Traffic Manager documentation, Azure WAF limits and rule set pages, the primary Microsoft and Cloudflare incident reports from October to December 2025, and AWS, Cloudflare and Google Cloud origin and WAF documentation, with calculated timing and original decision aids. Sources were reviewed on October 9, 2026.

No CDN, WAF, Traffic Manager profile or live traffic was configured or measured. Timing values are calculated from documented settings and exclude human decision time and resolvers that ignore TTLs. Incident details are the providers' own reports, and product limits are as documented on the review date.

AI assistance. AI assisted research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.

Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.

References

  1. High-availability implementation guide for using Azure Front Door and alternate ingress solutions Microsoft. Published . Accessed .
  2. Azure Traffic Manager endpoint monitoring Microsoft. Accessed .
  3. Cloudflare outage on November 18, 2025 Cloudflare. Published . Accessed .
  4. Cloudflare outage on December 5, 2025 Cloudflare. Published . Accessed .
  5. Secure traffic to Azure Front Door origins Microsoft. Accessed .
  6. Protect your origin server (Cloudflare Fundamentals) Cloudflare. Accessed .
  7. Proxy status (Cloudflare DNS) Cloudflare. Accessed .
  8. Validate JWTs (Cloudflare Access) Cloudflare. Accessed .
  9. Turnstile server-side validation Cloudflare. Accessed .
  10. Considerations for managing body inspection in AWS WAF Amazon Web Services. Accessed .
  11. Azure permissions for Networking (Azure RBAC) Microsoft. Accessed .
  12. Application Gateway WAF CRS and DRS rule groups and rules Microsoft. Accessed .