Research topic
Resilience
Recovery starts with an independent control path.
Design and testing guidance for containment, restoration, and control-plane recovery under pressure.
View every topicFrom the desk
26 publicationsRecover a deleted file with S3 Versioning
Recover an S3 object by inspecting retained versions, downloading a known copy, or removing the right delete marker without deleting the recovery data.
Choose an RDS recovery window you can actually restore
Choose RDS backup retention, inspect the real restorable interval, validate an isolated database restore, and record cleanup and recovery limits.
Recover deleted Google Cloud Storage objects with soft delete
Choose a Cloud Storage soft delete window, restore a specific object generation, validate its contents, and account for retention and cleanup.
Set up Cloud SQL backups and prove you can restore
Configure Cloud SQL PostgreSQL backups, understand retention settings, and rehearse a restore through database validation and application cutover.
Recover a deleted Azure blob with soft delete
Choose the right recovery action for a blob, version or container and verify restored bytes.
Recover a deleted Azure Key Vault secret
Distinguish secret recovery from vault recovery and verify versioned application references.
Make PostgreSQL point in time recovery reproducible
Build a version-aware recovery chain from protected base backups and WAL through timeline selection, isolated replay and application acceptance.
Plan Kubernetes drains around the disruption budget
Review selector scope, current status, unhealthy Pod handling and replacement capacity before treating a blocked drain as a reason to bypass availability controls.
Diagnose NAT gateway port exhaustion before adding capacity
Match allocation errors to destination tuples and gateway mode before changing connection pools, addresses or routes.
Separate stopping a fault experiment from recovering the service
Plan AWS FIS around separate evidence for stopping execution, removing fault effects and accepting the recovered application.
Find the EBS limit behind a slow database
Separate volume operation rate, byte rate, instance bandwidth and snapshot initialization before changing storage for a slow database.
Find the shared dependencies behind a cloud outage
Use the June 2025 Google Cloud and Cloudflare reports to review shared runtime, control, identity and recovery dependencies without turning one outage into a provider ranking.
Recovery objectives that match the cloud service
Define the business function, outage clock, recoverable data and dependency assumptions before choosing a cloud disaster-recovery architecture.
A controlled return from the SQS dead letter queue
Repair the failure, check consumer compatibility and return failed work with a bounded rate, observable stop conditions and business reconciliation.
Error budgets for controlled service degradation
Protect essential work under load while counting rejected and degraded requests against the service promise that users were actually given.
Certificate renewal under shorter validity limits
Use the public TLS issuance schedule to review authorization, renewal, deployment and independent verification of the certificate an endpoint actually serves.
The bottlenecks that shape a cloud DDoS response
Distinguish bandwidth, packet processing, connection state and application work before choosing a DDoS response or assuming the whole service path is protected.
Where DNS failover loses control of the clock
Separate authoritative routing, resolver caches, stale answers, runtime caching and existing connections when describing what DNS failover can achieve.
Keep database changes compatible with application rollback
Preserve an explicit relationship between old code and migrated state through additive changes, safe backfills, and a defined rollback window.
Stop retries from amplifying an outage
Count attempts across the complete request path, give retries a finite owner and budget, and define how repeated intent avoids duplicate side effects.
Keep encryption keys recoverable with the data they protect
Trace each encrypted recovery point to its required key, usable lifecycle state, and restore permissions before retiring cryptographic dependencies.
Make regional failover work without new infrastructure
Prepare capacity, dependencies, and the routing control path before an incident, then measure when clients reach an accepted recovery service.
Measure recovery by the service you can restore
Define application acceptance, recoverable data, and a complete timeline before treating a completed restore job as proof of recovery.
Protect backup copies from the account that runs production
Map deletion authority, retention protection, keys, and recovery identities so a surviving backup has a usable path back to service.
Questions answered
What does Cloud Security Desk cover under Resilience?
Design and testing guidance for containment, restoration, and control-plane recovery under pressure.
Supporting context
Recovery starts with an independent control path.
Why does the resilience topic matter to cloud security?
Recovery starts with an independent control path. This topic applies that principle to practical control evidence and review decisions.
Which control questions belong to Resilience?
Research belongs here when its central evidence, failure mode, or operational decision falls within the resilience boundary described on this page.
How many resilience publications are available?
This page currently lists 26 publications assigned to Resilience.
How should readers use the Resilience research?
Start with the publication closest to the control, platform, or evidence gap under review, then follow its findings, limitations, numbered sources, related work, and downloadable material.
Can resilience overlap other research topics?
Yes. A publication can appear under several topics when one evidence path crosses multiple control boundaries; each publication page identifies its complete scope.
Supporting context