
A threat analysis and detection guide for responders whose AWS, Azure or Google Cloud credentials are being used to run foundation models, built from provider documentation, Sysdig research and Microsoft's Storm-2139 update reviewed on October 9, 2026. It covers the calls attackers make, managed detections counted by provider, counting Bedrock calls and tokens without double counting, containment across both Bedrock endpoints, and rules to add.
At a glance
Key findings
- On the bedrock-runtime endpoint,
InvokeModel,InvokeModelWithResponseStream,ConverseandConverseStreamare CloudTrail management events, and every streamed call adds anAwsServiceEventthat carries token counts and the inference Region but must not be counted as an invocation. [1] - Inference on the bedrock-mantle endpoint is a CloudTrail data event that is off by default and outside model invocation logging, and the
AWSCompromisedKeyQuarantineV3policy edited March 16, 2026 names no bedrock-mantle action. [4][2][24] - GuardDuty AI Protection, a protection plan since July 13, 2026, baselines model, API, source and token volume per identity through a service-linked CloudTrail channel, and its three finding types default to Low severity. [11][12][13]
- Of 52 AI-specific detections that AWS, Microsoft and Google documented on October 9, 2026, 18 trigger on who reached the AI service, from where, by which method or how much. [13][11][14][15][16]
- Billing is the slowest signal: AWS Cost Anomaly Detection can take up to 24 hours to flag spend and needs 10 days of history for a service new to the account. [22]
How AI credential abuse shows up
When someone runs foundation models on your stolen credentials, the evidence is split across three kinds of record, and none of them is complete by itself. Control-plane logs say which identity called which model and from where: on Amazon Bedrock, InvokeModel, InvokeModelWithResponseStream, Converse and ConverseStream on the bedrock-runtime endpoint are CloudTrail management events, so an ordinary trail already holds them [1]. Token-bearing records say how much was consumed, but per call they exist only for streamed calls or only if someone configured them in advance: a second CloudTrail event written after each streamed response, and model invocation logs, which are off by default. CloudWatch metrics count tokens for every call but know the model, not the caller [1][2][3]. Billing says what it cost, a day or more later.
Start with identity and time window, not with the invoice. List every principal and access key that called a Bedrock runtime operation, count calls only from events whose eventType is AwsApiCall, take token counts from the stream-completion events and any invocation logs, and then look for what an attacker changes to stay hidden: a deleted invocation logging configuration, requests for model access, new Bedrock API keys, and traffic to the second inference endpoint, bedrock-mantle, whose inference calls are data events that a default trail never records [1][4].
The trigger is usually a cost alert, a GuardDuty finding, a production application suddenly throttled because someone else is spending its quota, or a provider's notice that a key was exposed. The cost alert is the slowest of these. Azure OpenAI and Google's Gemini Enterprise Agent Platform, formerly Vertex AI, split their evidence the same way under different names.
What the abuse pattern looks like
Most public detail about this attack comes from Sysdig's Threat Research Team, which coined the name LLMjacking, and from Microsoft's Digital Crimes Unit. Both describe incidents those organizations observed; they are the researchers' findings, not neutral measurements. In May 2024 Sysdig described credentials taken from a server running a vulnerable Laravel release (CVE-2021-3129) and a script that tested them against ten AI services, Bedrock and Vertex AI among them [5]. On Bedrock the checker ran no real prompt. It sent InvokeModel with max_tokens_to_sample set to -1, so a ValidationException confirmed model access where AccessDenied would have refused it. It also called GetModelInvocationLoggingConfiguration; Sysdig reports that OAI Reverse Proxy, the proxy whose user agent it saw, states that it will not use AWS keys with logging enabled [5].
Sysdig's September 2024 follow-up describes attackers enabling models instead of using only what was already on: ListFoundationModels and GetFoundationModelAvailability, then PutUseCaseForModelAccess and PutFoundationModelEntitlement, and in more recent intrusions DeleteModelInvocationLoggingConfiguration to stop prompt capture [6]. It counted more than 85,000 Bedrock requests over the period it monitored, 61,000 of them within three hours on July 11, 2024, and saw attackers adopt the Converse API within 30 days of its launch [6]. Its view of motive, from free personal use to resale to people in sanctioned countries, rests on prompts it could read precisely because those attackers never checked whether logging was on.
Microsoft's account is legal rather than technical. In an amended complaint described on February 27, 2025, it named four people at the center of a network it tracks as Storm-2139, which used exposed customer credentials scraped from public sources to reach generative AI services including Azure OpenAI, altered what those services would do and resold the access [7]. The sequence below condenses the Bedrock steps from both Sysdig reports into one path. No single incident is claimed to have followed every step, and the order of the middle steps varied.
The Bedrock abuse path in the 2024 research
Every step before resale is a call that CloudTrail records as a management event by default, which is where the hunt starts. [5][6][1]

Source. Conceptual composite of the Bedrock activity Sysdig reported in May and September 2024. [5][6]
Method. Conceptual sequence condensed from two Sysdig Threat Research Team reports; message names are the API calls those reports quote, except that the two logging-configuration calls, GetModelInvocationLoggingConfiguration and DeleteModelInvocationLoggingConfiguration, are described in words. No single incident is claimed to include every step, and the September report observed enablement and log deletion only in some intrusions.
Accessible table and figure data
| From | To | Message |
|---|---|---|
| Key checker | Bedrock runtime | InvokeModel with max tokens set to -1 |
| Bedrock runtime | Key checker | ValidationException confirms model access |
| Key checker | Bedrock control plane | Read the invocation logging configuration |
| Key checker | Bedrock control plane | ListFoundationModels, GetFoundationModelAvailability |
| Key checker | Bedrock control plane | PutUseCaseForModelAccess, PutFoundationModelEntitlement |
| Key checker | Bedrock control plane | Delete the invocation logging configuration |
| Reverse proxy | Bedrock runtime | Converse and InvokeModel calls for proxy customers |
| Bedrock runtime | Reverse proxy | Completions returned and resold |
| From | To | Message |
|---|---|---|
| Key checker | Bedrock runtime | InvokeModel with max tokens set to -1 |
| Bedrock runtime | Key checker | ValidationException confirms model access |
| Key checker | Bedrock control plane | Read the invocation logging configuration |
| Key checker | Bedrock control plane | ListFoundationModels, GetFoundationModelAvailability |
| Key checker | Bedrock control plane | PutUseCaseForModelAccess, PutFoundationModelEntitlement |
| Key checker | Bedrock control plane | Delete the invocation logging configuration |
| Reverse proxy | Bedrock runtime | Converse and InvokeModel calls for proxy customers |
| Bedrock runtime | Reverse proxy | Completions returned and resold |
The calls an attacker makes now
Several steps those reports keyed on have changed, and a rule set built only from them will miss part of the current path. Access to many Bedrock models is now enabled by default. The first invocation of a third-party model starts an AWS Marketplace subscription in the background, and calls can succeed during a setup period of up to 15 minutes [8]. It follows that an attacker need not call PutFoundationModelEntitlement for those models, and a rule that waits for it may never fire. The Bedrock page does not say which record the background subscription leaves, so record its absence from your logs as unknown rather than as proof that no new model was used.
Anthropic models still require a one-time use case submission per account or organization through PutUseCaseForModelAccess, but AWS states that the requirement does not apply when those models are reached through bedrock-mantle [8]. That endpoint serves OpenAI-compatible APIs and the Anthropic Messages API and authorizes inference with bedrock-mantle:CreateInference instead of bedrock:InvokeModel [9]. Its inference calls are CloudTrail data events, recorded only when a trail or event data store selects the AWS::BedrockMantle::Project resource type, they carry the event source bedrock-mantle.amazonaws.com, and model invocation logging does not capture them [4][2]. A detection or query that filters on bedrock.amazonaws.com sees none of it [4]. The bedrock-runtime endpoint has changed too: it now also accepts OpenAI-compatible Chat Completions and Responses calls and the Anthropic Messages API, authorized by bedrock:InvokeModel, while its CloudTrail page names events only for the four Bedrock-native operations [9][1]. Before trusting a filter on those four names, check which eventName values your trail actually holds for bedrock.amazonaws.com.
Bedrock API keys add a bearer-token path. A long-term key is a service-specific credential on an IAM user; one created in the console comes with a new user whose name starts BedrockAPIKey-, and its creation appears as CreateServiceSpecificCredential with serviceName set to bedrock.amazonaws.com [10]. A short-term key is generated on the client, so its creation never reaches CloudTrail, and it lives as long as the session that signed it, up to 12 hours [10]. Calls made with either kind carry callWithBearerToken set to true, in additionalEventData on bedrock-runtime events and in requestParameters on bedrock-mantle events [10][4].
Managed detections compared
All three providers now ship detections for AI services, and what each one reads, and how loudly it arrives, matters more than how many there are. On AWS, GuardDuty AI Protection became a protection plan on July 13, 2026, and its findings began feeding Extended Threat Detection attack sequences on September 17, 2026 [11]. Once enabled, it collects Bedrock, AgentCore and SageMaker AI data events through a CloudTrail service-linked channel that the account owner cannot reconfigure, so it does not depend on your trail [12]. Impact:IAMUser/AnomalousModelInvocation baselines the invocation API, model, source IP with its ASN organization, and user agent for each identity, and Impact:IAMUser/CostHarvesting baselines token volume; the third type, Impact:IAMUser/PromptInjection.Direct, depends on a Bedrock guardrail with a prompt attack filter [13]. All three default to Low severity, so a team that pages only on Medium and High will not see them without a routing change [13]. The foundational plan adds DefenseEvasion:IAMUser/BedrockLoggingDisabled, introduced on November 21, 2025 [11][12].
Microsoft Defender for Cloud's alert list for AI services, last updated July 6, 2026, includes a pair of wallet attack alerts, AI.Azure_DOWDuplicateRequests for runs of identical requests and AI.Azure_DOWVolumeAnomaly for request and response volume out of line with the resource's history, both Medium [14]. Beside them sit alerts for access from Tor and from addresses Microsoft threat intelligence flags, a suspicious user agent alert and AI.Azure_AccessAnomaly, which watches for changes in user agents, IP ranges and authentication methods against a resource [14].
Google's Security Command Center documents Event Threat Detection rules for Agent Platform assets on the Premium tier, and on the Enterprise tier until that tier shuts down on May 21, 2027 [15]. The three rules closest to this attack, New Geography for AI Service, New AI API Method and Dormant Service Account Activity in AI Service, are documented as reading Admin Activity audit logs [16]. Prediction calls such as endpoints.predict are Data Access operations, which are disabled by default [17]. It follows that, as documented, those three rules see the administrative calls around abuse rather than the inference traffic itself.
The chart counts every detection each provider documents as AI specific on the cited pages and splits off those whose documented trigger is who or what reached the service, from where, by which method or how much. Of 52 such detections, 18 are of that kind. The larger totals come mostly from agent, prompt and permission detections, which matter in other incidents but will not tell you that someone is reselling your model quota.
AI-specific managed detections by provider
Only 18 of 52 documented AI detections look at who reached the service or how much; GuardDuty has two, Defender seven and Security Command Center nine. [13][14][16][15]

Source. Counted on October 9, 2026 from GuardDuty AI Protection finding types and document history, Defender for Cloud alerts for AI services, and Security Command Center AI Protection and Event Threat Detection documentation. [13][11][14][15][16]
Method. Calculated count of named finding types, alerts or rules documented specifically for AI services. Identity or usage anomaly: the documented trigger is the identity, location, method or volume of access to the AI service. Excluded: runtime detector sets applied to agent hosts (Agent Platform Threat Detection, GuardDuty Lambda Protection) and generic rules not specific to AI. Google totals are the union of two pages that introduce their lists with include, so they are floors.
Accessible table and figure data
| Provider service | Identity or usage anomaly | Other AI detection |
|---|---|---|
| AWS GuardDuty | 2 | 2 |
| Microsoft Defender for Cloud | 7 | 11 |
| Google Security Command Center | 9 | 21 |
| Provider service | Identity or usage anomaly | Other AI detection |
|---|---|---|
| AWS GuardDuty | 2 | 2 |
| Microsoft Defender for Cloud | 7 | 11 |
| Google Security Command Center | 9 | 21 |
Logs that show which models ran
CloudTrail comes first because it exists by default and names the caller. Each runtime call carries the principal, the access key ID, the source IP, the user agent, awsRegion, the modelId in requestParameters and any errorCode [1]. A probe of the kind Sysdig described appears as InvokeModel with errorCode ValidationException [5]. If Bedrock rejected it at input validation, as that error suggests, it is also absent from the CloudWatch Invocations metric, which leaves out requests rejected before processing began [3].
Streamed calls need care. When InvokeModelWithResponseStream or ConverseStream finishes, Bedrock writes a second event with the same eventName, eventSource, userIdentity, source address and user agent, and eventType set to AwsServiceEvent [1]. It is not an API call and is never evaluated by IAM policies or service control policies, so counting it inflates volume and reading it as a request your deny failed to stop is wrong [1]. It is also the only CloudTrail record carrying inputTokens, outputTokens and inferenceRegion, under serviceEventDetails.AdditionalEventData.additionalEntries with a capitalized AdditionalEventData, and its serviceEventDetails.parentRequestId matches the call's requestID [1]. A trail that also records Bedrock data events gets a data-event copy of both, so one streamed call can produce four events; filter on eventCategory as well as eventType. inferenceRegion shows where a cross-Region inference profile ran the work, which can differ from the Region that logged it [1].
Non-streamed InvokeModel and Converse calls are the gap. The CloudTrail page documents token fields only for the stream-completion event, so take their usage from model invocation logs, which record identity.arn, modelId, operation, input.inputTokenCount and output.outputTokenCount per call, or from the InputTokenCount and OutputTokenCount metrics, whose only dimension is ModelId [2][3]. Invocation logging is off by default, covers only bedrock-runtime and keeps recording until the configuration is deleted [2]. That is why deletion is the attacker's move, and why the time of DeleteModelInvocationLoggingConfiguration marks where prompts and per-call usage stop [6]. For bedrock-mantle, the only per-call record is the CloudTrail data event, if your trail selected it. GuardDuty AI Protection reads data events through its own channel, and its documentation describes no copy delivered to you [12].
The fragment below applies those rules to an Athena table built with AWS's partition projection layout for CloudTrail. It counts calls from AwsApiCall events, joins token counts from stream-completion events through the parent request ID, and keeps the bearer-token flag so key use stands out. It does not cover bedrock-mantle; if you capture those data events, run a second query on bedrock-mantle.amazonaws.com and CreateInference. Widen the eventname list if your trail shows other runtime event names.
-- Calls come from AwsApiCall events; tokens from AwsServiceEvent stream-completion events.
-- Data-event copies are excluded so streamed calls are not counted twice.
WITH calls AS (
SELECT requestid,
useridentity.arn AS principal,
useridentity.accesskeyid AS access_key,
eventname, errorcode, sourceipaddress, useragent,
json_extract_scalar(requestparameters, '$.modelId') AS model_id,
json_extract_scalar(additionaleventdata, '$.callWithBearerToken') AS bearer_token
FROM cloudtrail_logs_pp
WHERE "timestamp" BETWEEN '2026/10/01' AND '2026/10/09'
AND eventsource = 'bedrock.amazonaws.com'
AND eventname IN ('InvokeModel', 'InvokeModelWithResponseStream', 'Converse', 'ConverseStream')
AND eventtype = 'AwsApiCall'
AND coalesce(eventcategory, '') <> 'Data'
),
usage AS (
SELECT json_extract_scalar(serviceeventdetails, '$.parentRequestId') AS parent_id,
CAST(json_extract_scalar(serviceeventdetails, '$.AdditionalEventData.additionalEntries.inputTokens') AS bigint) AS input_tokens,
CAST(json_extract_scalar(serviceeventdetails, '$.AdditionalEventData.additionalEntries.outputTokens') AS bigint) AS output_tokens,
json_extract_scalar(serviceeventdetails, '$.AdditionalEventData.additionalEntries.inferenceRegion') AS inference_region
FROM cloudtrail_logs_pp
WHERE "timestamp" BETWEEN '2026/10/01' AND '2026/10/09'
AND eventsource = 'bedrock.amazonaws.com'
AND eventtype = 'AwsServiceEvent'
AND coalesce(eventcategory, '') <> 'Data'
)
SELECT c.principal, c.access_key, c.model_id, c.bearer_token,
count(*) AS calls,
count_if(c.errorcode IS NOT NULL) AS failed_calls,
sum(u.input_tokens) AS streamed_input_tokens,
sum(u.output_tokens) AS streamed_output_tokens,
array_agg(DISTINCT u.inference_region) AS inference_regions
FROM calls c
LEFT JOIN usage u ON u.parent_id = c.requestid
GROUP BY c.principal, c.access_key, c.model_id, c.bearer_token
ORDER BY calls DESC;Parallel evidence in Azure and Google Cloud
On Azure, the control-plane record is the activity log, which keeps changes and initiated actions for 90 days and usually shows them within 3 to 20 minutes [18]. For an Azure OpenAI resource, pull Microsoft.CognitiveServices/accounts/listKeys/action and regenerateKey/action for who read or replaced keys, deployments/write for new model deployments, raiPolicies/write for content filtering changes, and Microsoft.Insights/DiagnosticSettings/Delete for removed logging [19][26]. Microsoft's description of Storm-2139 altering what services would do makes content filter changes worth a specific check [7]. Usage lives in metrics such as ProcessedPromptTokens, GeneratedTokens and TokenTransaction, split by deployment, model, API name and Region but not by caller, and in resource log categories such as RequestResponse and AzureOpenAIRequestUsage [20]. Resource logs are collected only once a diagnostic setting routes them, so they exist only if one was in place before the abuse [18].
Google Cloud draws the same line more sharply. Admin Activity audit logs always record configuration calls, but prediction calls are Data Access operations, which stay off until someone enables them for the Agent Platform service; a project that never enabled them has no per-call record of who ran which model [17]. Gemini API keys add a second gap. Google documents that a standard key ties requests to a project for billing and quota but does not identify a caller, while an authorization key, the default for keys created in Google AI Studio since May 28, 2026, runs requests as a bound service account [21]. Requests made with authorization keys are not counted in service account usage metrics, so a quiet service account in those metrics says nothing about its keys [21].
Where a cell in the matrix says a record exists only if it was enabled, the gap cannot be closed after the fact. Write it into the incident record as unknown, and size cost and notification decisions from what the remaining records support.
Where each investigative answer lives
Only the Bedrock runtime endpoint names the caller of each model call by default, and only its streamed calls carry token counts by default; per-call records elsewhere exist only if enabled in advance. [1][2][20][17]

Source. Source-derived summary of Bedrock CloudTrail, invocation logging and GuardDuty documentation, Azure OpenAI monitoring reference, Azure permissions and Defender alerts, and Agent Platform audit logging, Gemini API key and Event Threat Detection documentation, reviewed October 9, 2026. [1][4][2][13][10][20][19][26][14][17][21][16]
Method. Each cell restates a documented property in short form. Key calls unnamed is an inference from the absence of a caller dimension in Azure OpenAI metrics and from key-based authentication; the cited Google audit logging page lists prediction operations but no token fields.
Accessible table and figure data
| Question | Amazon Bedrock | Azure OpenAI | Google Agent Platform |
|---|---|---|---|
| Which identity called | CloudTrail runtime events, default | Activity log for key reads; key calls unnamed | Data Access logs, only if enabled |
| Which model ran | modelId in CloudTrail; mantle needs data events | Metrics by deployment and model | Data Access logs, only if enabled |
| How many tokens | Stream events; invocation logs if enabled | Token metrics; resource logs if routed | Not in audit log documentation |
| Logging tampered | CloudTrail delete of logging configuration | Diagnostic setting Delete action | Data Access logging is opt-in |
| Keys created or read | Service-specific credential creation | listKeys and regenerateKey actions | Gemini standard keys name no caller |
| Managed detection | GuardDuty AI Protection, Low | Defender wallet and access alerts | ETD AI rules on Admin Activity |
| Question | Amazon Bedrock | Azure OpenAI | Google Agent Platform |
|---|---|---|---|
| Which identity called | CloudTrail runtime events, default | Activity log for key reads; key calls unnamed | Data Access logs, only if enabled |
| Which model ran | modelId in CloudTrail; mantle needs data events | Metrics by deployment and model | Data Access logs, only if enabled |
| How many tokens | Stream events; invocation logs if enabled | Token metrics; resource logs if routed | Not in audit log documentation |
| Logging tampered | CloudTrail delete of logging configuration | Diagnostic setting Delete action | Data Access logging is opt-in |
| Keys created or read | Service-specific credential creation | listKeys and regenerateKey actions | Gemini standard keys name no caller |
| Managed detection | GuardDuty AI Protection, Low | Defender wallet and access alerts | ETD AI rules on Admin Activity |
Cost signals and their delay
Billing is the slowest and least specific signal, and the providers say so. AWS Cost Anomaly Detection runs about three times a day on Cost Explorer data that can lag by up to 24 hours, and a service new to the account needs 10 days of usage history before anomalies can be detected for it [22]. A reasonable reading for an account that never used Bedrock is that the first ten days of abuse fall outside anomaly detection for that service. GuardDuty's CostHarvesting finding works from token volume in data events instead, but it baselines against each identity's own history and arrives at Low severity [13].
Google added two controls in July 2026 that narrow the delay. Early anomalies for AI workloads such as Gemini API and Vertex AI, released July 24, use near real-time cost estimates, and spend cap budgets, in preview since July 27, pause new requests to an eligible service in a project once estimated costs pass the budget [23]. Google notes that enforcement is faster than billing reports but not instant, and overages are still billed [23]. Because a spend cap also stops legitimate traffic, treat it as containment for a project you can afford to pause.
For reconciliation, compute exposure from your own token counts, not from published estimates. Sysdig's often repeated figure of more than USD 46,000 a day was its 2024 worst case for Claude 2, assuming 500,000 tokens a minute in four Regions at an averaged USD 0.016 per 1,000 tokens [5]. Models, quotas and prices have changed since. Price the tokens by model and inference Region from the queries above against the list in force during the window, and use the bill as a cross-check once it settles.
| Signal | Documented timing or basis | Limit |
|---|---|---|
| AWS Cost Explorer data | Up to 24 hours behind usage | Feeds anomaly detection |
| AWS Cost Anomaly Detection | About three runs a day | 10 days of history for a new service |
| GuardDuty CostHarvesting | Token volume against a baseline | Default severity Low |
| Google early anomalies for AI | Near real-time cost estimates | User thresholds do not apply |
| Google spend cap budget | Estimated cost passes the budget | Preview; not instant; pauses new requests |
Contain without losing evidence
Capture before you change anything. Record the output of GetModelInvocationLoggingConfiguration in every Region where Bedrock is used, note whether the trail selects Bedrock and bedrock-mantle data events and whether GuardDuty AI Protection is enabled, and copy the CloudTrail window for the affected principals and keys out of any store whose retention could expire it. Then contain the AI path, which means both endpoints and both credential types.
AWS's managed AWSCompromisedKeyQuarantineV3, last edited March 16, 2026, denies bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:CreateModelInvocationJob, bedrock:CreateFoundationModelAgreement and bedrock:PutFoundationModelEntitlement [24]. Its action list names no bedrock-mantle action and neither CallWithBearerToken action, so by its text it does not deny inference through bedrock-mantle; whether a quarantined user can still reach that endpoint depends on the user's other permissions [24][9]. Because a call made with a Bedrock API key is authorized against the IAM principal behind it, a deny on the principal also stops key-based calls [9]. The fragment at the end of this section adds the missing actions. Attach it to the user or role under investigation, not to a shared role production depends on, unless stopping that traffic is the decision you have made.
Then handle the credentials. Deactivate a long-term Bedrock API key with aws iam update-service-specific-credential and --status Inactive before deleting anything, so the credential ID stays available for the record [10]. For a short-term key, AWS's guidance is to revoke the IAM role session that produced it or deny bedrock:CallWithBearerToken and bedrock-mantle:CallWithBearerToken to the identity [10]. Restore invocation logging only after the deny is in place, so the attacker cannot simply delete it again, and keep the deny until a controlled call with the old credential fails with an authorization error.
On Azure, setting disableLocalAuth to true changes the control plane at once, but Microsoft documents that the gateway can keep accepting previously valid keys until its cached configuration refreshes, usually within minutes and sometimes after several hours, and tells administrators to plan for that delay when they rotate keys as well [25]. Confirm the cutoff with a request using the old key that returns HTTP 401. On Google, disable the exposed API key or service account key, then enable Data Access audit logs for Agent Platform if they were off, so the next request leaves a record [17].
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyInferenceOnBothEndpoints",
"Effect": "Deny",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream",
"bedrock:CreateModelInvocationJob",
"bedrock-mantle:CreateInference",
"bedrock:CallWithBearerToken",
"bedrock-mantle:CallWithBearerToken"
],
"Resource": "*"
},
{
"Sid": "DenyLoggingAndAccessChanges",
"Effect": "Deny",
"Action": [
"bedrock:DeleteModelInvocationLoggingConfiguration",
"bedrock:PutModelInvocationLoggingConfiguration",
"bedrock:PutUseCaseForModelAccess",
"bedrock:CreateFoundationModelAgreement",
"iam:CreateServiceSpecificCredential"
],
"Resource": "*"
}
]
}Rules to write yourself
Most managed findings that bear on this attack default to Low or Medium severity and cover only part of the path, so a handful of rules of your own close the gaps that matter here. Most key on rare control-plane events or first-seen combinations, so they stay quiet where a small team owns model use.
- Any
DeleteModelInvocationLoggingConfigurationorPutModelInvocationLoggingConfigurationoutside an approved change, in any Region. GuardDuty's foundational plan detects disabled invocation logging; a rule of your own also catches aPutthat changes what is delivered without turning logging off [12][2]. InvokeModelorConversefailing withValidationExceptionorAccessDeniedExceptionfrom a principal with no successful runtime call in the previous 30 days, the probe shape Sysdig described [5].PutUseCaseForModelAccessorCreateFoundationModelAgreementfrom anyone outside the platform team, plus a first-seenmodelIdper principal, because default access means many first uses leave no access request at all [8].CreateServiceSpecificCredentialwithserviceNamebedrock.amazonaws.com,CreateUserfor a name beginningBedrockAPIKey-, and runtime calls withcallWithBearerTokentrue from principals with no approved key use [10].- Any event from
bedrock-mantle.amazonaws.comin an account that does not use that endpoint, and a posture check that trails select mantle data events wherever it is used [4]. - A first-seen
awsRegionorinferenceRegionper principal; Sysdig saw its attackers sendInvokeModelonly to the Regions where Bedrock was offered [5][1]. - On Azure,
listKeys/actionby an identity outside the deployment pipeline, andraiPolicies/writeorDiagnosticSettings/Deleteagainst Cognitive Services accounts [19][26]. - On Google, Data Access audit logs enabled for Agent Platform in every project that calls models, with an alert on prediction calls from service accounts outside each project's usual set [17].
Order of work for the first day
One rule governs scope until the evidence narrows it: if you cannot yet name the identity, every key, session and API key that can call Bedrock, Azure OpenAI or Agent Platform in the affected account is in scope. Narrow from the logs, not from the bill. Then work in this order, and do not move to the next step until its check passes.
- Capture: logging configuration per Region, trail export for the window, data-event selection and GuardDuty AI Protection status. Check: you can say which records exist and which never did.
- Scope: calls by principal, key and model from
AwsApiCallevents, tokens from stream-completion events and invocation logs, and mantle data events if captured. Check: per-model call and token totals roughly agree with the CloudWatch metrics for the same window. - Contain: deny inference on both endpoints and bearer tokens for the principal, deactivate keys, end sessions. Check: a controlled call with the old credential fails with an authorization error, and on Azure the old key returns 401.
- Restore: invocation logging, Azure diagnostic settings or Google Data Access logs, plus the rules above. Check: the next streamed call produces both CloudTrail events and an invocation log entry.
- Reconcile: price the token counts once billing settles, and keep the gaps you recorded as unknown in the final report.
Method and provenance
Source-led threat analysis of AWS, Microsoft and Google Cloud documentation, Sysdig threat research and a Microsoft Digital Crimes Unit legal update, with detection counts calculated from provider pages. Sources were reviewed on October 9, 2026, and the detection counts, Bedrock event names and data-event statements were re-checked against the same pages on October 10, 2026.
No cloud account, model or log pipeline was used, and no query or policy was run. Detection counts reflect the cited pages on the review date and change often; threat research describes incidents those researchers observed and is not a measure of prevalence.
AI assistance. AI assisted research synthesis, drafting, query and policy examples, chart and diagram planning, and visual production. No personal experience or live test is claimed; the publication's editors are responsible for the content.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Monitor Amazon Bedrock API calls using CloudTrail Amazon Web Services. Accessed .
- Monitor model invocation using CloudWatch Logs and Amazon S3 Amazon Web Services. Accessed .
- Monitor bedrock-runtime inference using CloudWatch metrics Amazon Web Services. Accessed .
- Monitor bedrock-mantle API calls using CloudTrail Amazon Web Services. Accessed .
- LLMjacking: Stolen Cloud Credentials Used in New AI Attack Sysdig. Published . Accessed .
- The Growing Dangers of LLMjacking: Evolving Tactics and Evading Sanctions Sysdig. Published . Accessed .
- Disrupting a global cybercrime network abusing generative AI Microsoft. Published . Accessed .
- Request access to models (Amazon Bedrock User Guide) Amazon Web Services. Accessed .
- Endpoints supported by Amazon Bedrock Amazon Web Services. Accessed .
- Securing Amazon Bedrock API keys: Best practices for implementation and management (updated July 1, 2026) Amazon Web Services. Published . Accessed .
- Document history for Amazon GuardDuty Amazon Web Services. Accessed .
- GuardDuty AI Protection Amazon Web Services. Accessed .
- GuardDuty AI Protection finding types Amazon Web Services. Accessed .
- Alerts for AI services (Microsoft Defender for Cloud) Microsoft. Published . Accessed .
- AI Protection overview (Security Command Center) Google Cloud. Accessed .
- Event Threat Detection overview Google Cloud. Accessed .
- Agent Platform audit logging information Google Cloud. Accessed .
- Activity log in Azure Monitor Microsoft. Published . Accessed .
- Azure permissions for AI + machine learning Microsoft. Published . Accessed .
- Azure OpenAI monitoring data reference Microsoft. Published . Accessed .
- Using Gemini API keys Google. Accessed .
- Detecting unusual spend with AWS Cost Anomaly Detection Amazon Web Services. Accessed .
- Cloud Billing release notes Google Cloud. Accessed .
- AWSCompromisedKeyQuarantineV3 (AWS managed policy reference) Amazon Web Services. Accessed .
- Disable local authentication in Foundry Tools Microsoft. Published . Accessed .
- Azure permissions for Monitor Microsoft. Published . Accessed .