
An architecture analysis for AI platform and API gateway engineers, based on Azure API Management, Apigee, RedisVL and GPTCache documentation and source and three peer-reviewed attack papers reviewed on October 10, 2026. It follows a request through the cache, shows where documented defaults lose the caller's scope, and gives a partition design, a threshold comparison, a policy fragment and a two-user test plan.
At a glance
Key findings
- API Management's own semantic caching examples partition entries by
context.Subscription.Id, a key container that the subscriptions documentation ties to applications and to teams sharing keys, while its usage notes call for user or user-group identifiers invary-byto control cross-user access. [1][2][3] - RedisVL searches the whole cache unless each
check()call passes a filter expression and keeps entries forever at its defaultttl=None; Apigee's lookup policy lists no partition element at all. [6][7][10] - Threshold direction differs by product: lower is stricter for API Management
score-thresholdand RedisVLdistance_threshold, higher is stricter for GPTCache and for Apigee with dot product. [1][7][9][10] - In NDSS 2026 black-box tests on deployments the authors built on AWS, Azure and Alibaba, ordinary users poisoned the semantic cache in 76 to 93 percent of attempts, and in a GPTCache sweep success held until the threshold passed 0.95. [14]
- A key collision attack reached an 86.7 percent hit rate against Azure API Management semantic caching using only a public surrogate embedding model; a secret key salt cut hit rates by at most 21.0 points. [15]
A cache hit is an authorization decision
A semantic response cache answers a narrower question than the model does: has anyone in this partition already asked something close enough to this prompt. When the answer is yes, it returns the stored response without calling the model, without rerunning retrieval or tools, and without checking anything about the current caller. Vector similarity has no notion of permission, so authorization has to be settled by which partition the lookup may search. Microsoft's reference for the API Management llm-semantic-cache-lookup policy says this in its usage notes, which tell readers to control cross-user access with vary-by set to user or user-group identifiers, and it warns that similarity-based reuse can return responses that are incorrect, outdated or unsafe for the current request. [1]
The design answer is an exact-match partition, built from trusted values before any similarity runs. It names the tenant; the end user, or a fingerprint of the entitlements that shaped the answer; the calling application and its prompt template version; the model deployment; and, where retrieval or tools contributed, either those inputs or a scope that bounds them. An answer that depended on one person's data is stored only in that person's partition, or not at all. Inside a partition, a strict threshold calibrated on labeled prompt pairs keeps honest false hits rare, and a store path that only the gateway or orchestrator writes limits what a user can plant for others.
Published research sets the limit on that design. In NDSS 2026 experiments, ordinary users poisoned semantic caches that the authors deployed on three clouds in 76 to 93 percent of black-box attempts. A key collision method reached an 86.7 percent hit rate against Azure API Management semantic caching, and a timing attack told a neighbor's cached prompt from a miss 81.4 percent of the time on a single try. [14][15][16] The size and trust level of the population sharing a partition is the security parameter. The similarity threshold is mainly a correctness parameter.
Follow one request through the cache
The gateway products share one path. A request arrives, the gateway extracts the text it will compare, calls an embeddings model, searches a vector index for the nearest stored prompt, and compares the score with a threshold. A hit returns the stored response. A miss goes to the model, and the response path stores the prompt vector and the response with an expiry. In API Management the lookup runs in the inbound section and calls embeddings through a backend named by embeddings-backend-id using the instance's system-assigned identity; the store policy runs in the outbound section; the cache itself is an external Azure Managed Redis instance whose RediSearch module can only be enabled when the cache is created. [1][2] Apigee's SemanticCacheLookup policy uses Vertex AI text embeddings and Vector Search instead. [10]
What gets embedded decides what the key can tell apart. API Management's ignore-system-messages, which Microsoft recommends setting to true, removes system messages before similarity is assessed, and max-message-count skips caching once the remaining dialog passes a set number of messages. [1] Apigee's default UserPromptSource reads $.contents[-1].parts[-1].text, the last part of the last message. [10] Each setting is sensible for hit rate. Between them, they remove from the comparison the parts of a request most likely to carry authorization context: system instructions that name a role or tenant, earlier turns, and retrieved passages that the application places in a system message.
Microsoft's Cosmos DB guidance on semantic caching shows the correctness version of the problem. One user asks for the largest lake in North America and follows with "What is the second largest?"; another asks about the largest stadium and sends the same follow-up, and a cache keyed on the last prompt alone answers the stadium question with Lake Huron. The guidance concludes that a semantic cache should operate within the conversation's context window, and it describes the returned completion as one that another user's earlier sequence produced. [12] The authorization version is the same mechanism with a permission boundary where the topic boundary was.
The store side carries the same blind spot. The response stored on a miss is whatever the model produced for that caller, using that caller's documents and tool results, and the entry records none of those inputs unless the application adds them. A RedisVL entry holds the prompt, the response, the vector, timestamps and optional metadata and filters. The AWS sample read-through cache passes its Lambda cache handler only a prompt, a generation length and a reset flag. [6][7][13]
Where authorization falls out of the key
Consider a hypothetical internal assistant behind API Management. Every employee uses the same web application, which calls the gateway with one subscription key. A finance manager asks for the approved salary bands for her team; the application retrieves compensation documents she is allowed to read and the model summarizes them. Later a sales lead asks for his team's salary bands. The two prompts embed close together, the partition is the subscription, and the sales lead receives the finance answer. The retrieval-time permission check never ran for him, because a hit skips retrieval.
Microsoft's examples partition by context.Subscription.Id in both the policy reference and the how-to. [1][2] An API Management subscription is a named container for a pair of keys. The subscriptions documentation describes each application rotating between its two keys, describes standalone subscriptions shared by several developers or teams, and notes that subscriptions cannot be assigned to Microsoft Entra ID security groups. [3] So the example's partition is normally an application or a team, not a person. That fits a single-purpose endpoint that answers from public material. It is the wrong boundary for anything that reads per-user data, and the replacement value should come from a token the gateway has validated, not from a header the client controls. [4]
RedisVL makes scoping a per-call choice. filterable_fields defines tag fields such as user_id, but check() applies them only when the caller passes a filter_expression, and the v0.28.0 source documents that without one the full cache is searched. [6][7] The library guide's own multi-user example stores two users' phone numbers in one index and relies on every lookup carrying the user filter, so a single code path that forgets the filter becomes a cross-user read. A reasonable reading is that the filter belongs in a wrapper that refuses to look up without a scope, or that each tenant gets its own index name.
Apigee's reference for SemanticCacheLookup, updated October 7, 2026, lists prompt source, embeddings and similarity search elements, and no element for a partition key or a Vector Search filter. Its URLs accept templating, which is one way to send each scope to its own deployed index. [10] Google's tutorial recommends not caching user-specific responses, or giving them a short TTL such as five minutes. [11] A short TTL limits how long an entry is exposed. It does not change who can hit the entry while it lives.
What the published attacks measured
Three peer-reviewed studies matter for the design, and each describes the attacker as an ordinary user who can send prompts and observe only their own responses and timing. Wu and colleagues, at NDSS 2026, poison the cache. The attacker writes a prompt that the embedding model scores as close to a target question but that steers the model, often with plain instructions rather than injection strings, toward an attacker-chosen answer. That pair is stored, and later users who ask the target question or a paraphrase receive the planted answer; in the black-box setting 95.7 percent of paraphrases of a poisoned target also returned it. [14]
The cloud results come from deployments the authors built themselves, black-box, under default or recommended settings. The paper labels the AWS deployment AWS Bedrock and describes the cloud caches as built-in features, but the source it cites is the 2024 AWS sample read-through cache built from OpenSearch Serverless, Bedrock and a Lambda handler, run at a 0.75 cosine threshold; this article follows the cited source and treats it as a sample architecture, not a managed Bedrock feature. The Azure deployment is cited to Microsoft's Cosmos DB semantic cache guidance with a 0.8 threshold, and Alibaba Higress ran with 0.8 applied; all three used a 10-second expiry. [12][13][14] Prompt-level defenses based on perplexity, paraphrasing and injection classifiers averaged F1 scores of 0.37 to 0.53. Their own check, which scores the retrieved response against the new query, reached 0.80 to 0.92, and the authors say it cannot remove the risk, particularly for subjective questions. They report that Alibaba confirmed the issue, while AWS and Azure were still investigating when the paper was written. [14]
Zhang and colleagues, at ICML 2026, treat the semantic key as a fuzzy hash that by design lacks the avalanche property of a cryptographic one. Their CacheAttack-2 method tunes a suffix against a public surrogate embedding model and confirms each candidate with one query to the target. It reached hit rates of 83.1 percent on GPTCache at a 0.8 threshold, 78.2 percent on the AWS sample and 86.7 percent on Azure API Management, against 12.4 and 15.2 percent for a genetic-algorithm baseline on the two cloud setups; the paper does not detail the gateway configuration. When cache entries held tool invocations, the correct tool selection rate fell by 84.5 points. A secret salt mixed into the key input cut hit rate by at most 21.0 points, and per-user namespaces removed cross-user hijacking at a cost in hit rate. [15]
Song and colleagues, in a paper accepted by IEEE Transactions on Information Forensics and Security, read latency instead of content. Under GPTCache defaults a hit returned in under a second against about five seconds for a miss, enough to tell 81.4 percent of the time on one try, and 95.4 percent after five, whether a neighbor had sent a prompt containing particular private attributes. They suggest anonymizing private attributes before the similarity search, at about 4 percent overhead. [16] The OWASP Top 10 for LLM Applications 2026 now lists semantic cache poisoning under LLM09:2026 Vector and Embedding Weaknesses. [17]
Read these numbers as feasibility under default or sample configurations, not as prevalence. Their design consequence is narrower and firmer: within a partition shared by people who do not trust each other, a strict threshold does not stop a determined insider, so the partition boundary and the write path have to do that work.
Semantic cache poisoning success in black-box tests, NDSS 2026
Every deployment and prompt method succeeded in at least three of four attempts, using thresholds of 0.75 to 0.8. [14]

Source. Wu et al., NDSS 2026, Table III, black-box columns; thresholds and setup from Section VI-A. Deployments built by the authors under default or recommended settings with a 10-second expiry; success judged by an LLM judge. Reviewed October 10, 2026. [14]
Method. Values copied from Table III without transformation. The GPTCache white-box and text-to-image columns are omitted so that every bar is a black-box result. The AWS row is the AWS sample architecture the paper cites, not a managed Bedrock feature. [13][14]
Accessible table and figure data
| Deployment | Zero-shot prompting | In-context learning | Prompt injection templates |
|---|---|---|---|
| AWS sample (OpenSearch Serverless and Bedrock), 0.75 | 79 | 76 | 87 |
| Alibaba Higress, 0.8 applied | 92 | 77 | 93 |
| Azure (Cosmos DB guidance), 0.8 | 85 | 90 | 91 |
| GPTCache text-to-text, black-box, 0.8 | 84 | 78 | 94 |
| Deployment | Zero-shot prompting | In-context learning | Prompt injection templates |
|---|---|---|---|
| AWS sample (OpenSearch Serverless and Bedrock), 0.75 | 79 | 76 | 87 |
| Alibaba Higress, 0.8 applied | 92 | 77 | 93 |
| Azure (Cosmos DB guidance), 0.8 | 85 | 90 | 91 |
| GPTCache text-to-text, black-box, 0.8 | 84 | 78 | 94 |
Build the partition before similarity
Treat the partition as a tuple of exact-match fields resolved at a trusted boundary, and similarity as a search inside one tuple. The fields below are an original synthesis of the documented behavior and the attack results, not a vendor checklist. Each one either keeps an answer inside the population allowed to receive it or keeps it tied to the inputs that made it correct, and the figure shows where each is settled on the request path.
- Tenant: taken from the validated token, never from the request body.
- Authorization scope: the end user for answers built from personal data or per-user tools; otherwise a hash of the entitlement set that the authorization service resolved, so people with identical access share and an access change moves a user to a new partition.
- Application and prompt template version: required because the recommended
ignore-system-messagessetting removes system text from the similarity input. [1] - Model deployment and version: a correctness field rather than an authorization field, kept in the key because a model change changes answers.
- Retrieval and tool inputs: a gateway cache runs before retrieval and cannot see which documents will be used, so either bound them with the authorization scope or move the lookup after retrieval and key on the retrieved document identifiers and versions.
A cache path that settles scope before similarity
Similarity search runs only inside an exact-match partition built from validated claims, and only shareable responses are written back. [1][4][15]

Source. Conceptual architecture based on API Management semantic caching and token validation documentation and the per-user isolation discussion in the CacheAttack paper. [1][4][15]
Method. Conceptual. Edges are authored to show the order of checks; this is a design proposal, not a measured deployment.
Accessible table and figure data
| Component | Role |
|---|---|
| Verified caller | Token validated at the gateway before any cache step |
| Scope resolver | Turns claims into tenant and user or entitlement hash |
| Exact-match partition | Scope plus application, template and model |
| Similarity search | Nearest neighbor inside one partition only |
| Cached response | Returned on a hit, never across partitions |
| Model call | Runs retrieval and tools as the current caller |
| Store decision | Writes shareable classes back to the same partition |
| Component | Role |
|---|---|
| Verified caller | Token validated at the gateway before any cache step |
| Scope resolver | Turns claims into tenant and user or entitlement hash |
| Exact-match partition | Scope plus application, template and model |
| Similarity search | Nearest neighbor inside one partition only |
| Cached response | Returned on a hit, never across partitions |
| Model call | Runs retrieval and tools as the current caller |
| Store decision | Writes shareable classes back to the same partition |
Treat the store as a write path
Responses produced with per-user tools, personal records or write actions go into the user's own partition or are not stored, and cached tool invocations are a poor fit for any shared tier given the CacheAttack results. Make the gateway or orchestrator the only writer, record the policy version that admitted each entry, and fill any shared tier for common questions from reviewed content rather than from whatever users happen to ask. The NDSS authors argue that a semantic cache needs cross-user sharing to be worth running, which rules out isolating everyone. Both integrity attacks, as published, depend on the attacker's own query becoming a stored entry, so a reasonable inference is that a shared tier users cannot write into is the form of sharing they do not reach; if that tier holds only reviewed public content, a timing hit on it also reveals nothing private. [14][15]
The fragment shows the API Management side of a per-user tier. It validates the Entra token into a variable, rejects tokens without an object ID rather than letting an empty value become a shared partition, partitions by tenant, user, subscription and prompt version, and rate limits per user after the lookup as Microsoft recommends. [1][4][5] The same structure works for an entitlement-hash tier if the hash arrives as a claim in a token the gateway validates, not in an ordinary header.
<!-- Example fragment: per-user semantic cache tier in an API-scope policy. -->
<inbound>
<base />
<validate-azure-ad-token tenant-id="{{entra-tenant-id}}" output-token-variable-name="jwt">
<client-application-ids>
<application-id>{{assistant-client-app-id}}</application-id>
</client-application-ids>
</validate-azure-ad-token>
<!-- Reject instead of letting an empty claim become one shared partition. -->
<choose>
<when condition="@(string.IsNullOrEmpty(((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("oid", "")))">
<return-response>
<set-status code="403" reason="Forbidden" />
</return-response>
</when>
</choose>
<llm-semantic-cache-lookup
score-threshold="0.05"
embeddings-backend-id="embeddings-backend"
embeddings-backend-auth="system-assigned"
ignore-system-messages="true">
<vary-by>@(((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("tid", ""))</vary-by>
<vary-by>@(((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("oid", ""))</vary-by>
<vary-by>@(context.Subscription.Id)</vary-by>
<vary-by>@(context.Request.Headers.GetValueOrDefault("x-prompt-version", "none"))</vary-by>
</llm-semantic-cache-lookup>
<rate-limit-by-key calls="20" renewal-period="60"
counter-key="@(((Jwt)context.Variables["jwt"]).Claims.GetValueOrDefault("oid", ""))" />
</inbound>
<outbound>
<base />
<llm-semantic-cache-store duration="300" />
</outbound>Thresholds narrow false hits, not attacks
The four products express the threshold on different scales and in opposite directions, so a value copied between them can mean the reverse of what was intended. API Management's score-threshold runs from 0.0 to 1.0 and lower values demand closer matches; the reference example uses 0.05, the how-to uses 0.15, and the usage notes warn that values above 0.2 may lead to mismatches. [1][2] RedisVL's distance_threshold is a Redis cosine distance from 0 to 2 with a default of 0.1, and its guide shows that at 0.5 the question about the capital of the country where Nice is located returns the cached answer Paris. [6][7]
GPTCache's similarity_threshold defaults to 0.8, and adapter.py counts a hit when the evaluator's score reaches the threshold times the evaluator's range, so higher is stricter. The parameter's docstring in config.py describes the two endpoints the other way round, which is a reason to read the comparison in code before tuning. [8][9] Apigee's <Threshold> defaults to 0.9, a value meant for DOT_PRODUCT_DISTANCE, where a match must reach the threshold; for cosine, squared L2 and L1 distance a match must fall at or below it. Apigee accepts and silently ignores <DistanceMeasureType> before version 1-18-0-apigee-4 and hybrid 1.17.0, comparing as dot product. [10]
Calibrate on labeled pairs for the embedding model in use: paraphrases that should share an answer, and near misses that must not, such as a negation, a different entity, a different date or a different team. Choose the loosest threshold whose false-hit rate on the near misses is acceptable for the use case, rather than the one that maximizes hit rate. The RedisVL guide makes the same point that the right value depends on the embedding model, the queries and the use case. [6]
Do not expect the threshold to stop an attacker. NDSS adversarial prompts reached an average similarity of 0.87 to their targets in the black-box setting, and in the GPTCache sweep attack success stayed level until the threshold exceeded 0.95. CacheAttack's hit rate declined only gradually as its threshold rose from 0.75 to 0.90. [14][15] The NDSS authors note that a threshold high enough to break the attack also makes the system less usable. [14]
| Product and setting | Scale and stricter direction | Default or documented value |
|---|---|---|
| API Management score-threshold | Score 0.0 to 1.0; lower is stricter | Required, no default; examples 0.05 and 0.15; above 0.2 may mismatch |
| RedisVL distance_threshold | Cosine distance 0 to 2; lower is stricter | 0.1 |
| GPTCache similarity_threshold | Share of evaluator range 0 to 1; higher is stricter | 0.8 |
| Apigee Threshold with dot product | Similarity; higher is stricter | 0.9; tutorial suggests 0.95 |
| Apigee Threshold with cosine, L1 or squared L2 | Distance; lower is stricter | Reset it: the 0.9 default is meant for dot product |
Expiry, invalidation and deletion
Expiry defaults differ as much as thresholds. API Management makes the store duration a required number of seconds, and both Microsoft examples use 60. [1][2][18] Apigee's populate policy defaults to 60 seconds and ignores cache-control headers from the model. [11] RedisVL's default ttl=None keeps entries forever. When a TTL is set, every check() hit refreshes it on all matched entries, so a popular entry can live indefinitely, and setting a TTL later adds one only to entries that are matched again. [6][7]
Revocation is where partition design and expiry meet. If the partition is the user's object ID, an answer built from documents the user could once read stays reachable by that user after access is withdrawn, until the entry expires. If the partition includes a hash of the resolved entitlement set, the access change moves the user to a new partition at once and the old entry becomes unreachable, although it still exists. That is a strong reason to prefer an entitlement hash wherever the authorization service can supply one cheaply.
Unreachable is not deleted. Cached prompts, responses and prompt vectors are copies of user content, and the vectors are derived data that deserve the same classification. The two API Management semantic caching policy pages describe only lookup and store attributes, so purge and invalidation have to be planned at the external Redis instance; RedisVL exposes drop() by entry ID or key, clear() and delete(); Google points Apigee users to Vertex AI data point updates. [1][18][6][11] Keep the partition fields readable in each entry so that one user's entries, one tenant's entries or one suspected poisoned entry can be found and removed without flushing everything.
Test the cache with two users
Run these checks in an authorized test environment with synthetic users and fixtures. Each one has an expected outcome that can be stated before it runs. API Management shows cache use in a request trace, Apigee adds a Cached-Content: true response header and exposes a cache_hit flow variable, and a RedisVL lookup returns its matched entries directly. [2][10][11]
The negative cases only mean something after the positive control passes. If the same user's paraphrase does not hit, the cache is not working, and every miss below proves nothing. The two-principal discipline is the same one that row-level security tests in a database depend on.
- Positive control: one user sends a prompt and then a paraphrase; the second request hits.
- Same tenant, different entitlements: user B sends user A's exact prompt and misses.
- Different tenants: the same prompt from another tenant misses.
- Missing identity: a token without the object ID claim is rejected, not served from an empty partition.
- Version change: a new prompt template or model deployment misses on the first request.
- Near misses: negation, entity and date variants miss at the chosen threshold.
- Shared tier writes: a user's crafted answer to a common question never reaches another user's lookup.
- Revocation and expiry: after access is withdrawn or the TTL passes, the earlier answer is not returned.
Method and provenance
Source-led architecture analysis of Microsoft, Google Cloud, Redis and GPTCache documentation and source code, an AWS sample architecture, and three peer-reviewed attack papers, with original diagrams. Sources were reviewed on October 10, 2026.
No gateway, cache or model was configured or tested. Attack figures come from the authors' own deployments under default or sample settings and describe feasibility, not prevalence. Product behavior is bounded to the cited pages and source versions as of the review date.
AI assistance. AI assisted research synthesis, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Enable semantic caching for LLM APIs in Azure API Management Microsoft. Accessed .
- Subscriptions in Azure API Management Microsoft. Accessed .
- Validate Microsoft Entra token (validate-azure-ad-token policy) Microsoft. Accessed .
- API Management policy expressions Microsoft. Accessed .
- Cache LLM Responses (RedisVL user guide) Redis. Accessed .
- redisvl/extensions/cache/llm/semantic.py at v0.28.0 Redis. Accessed .
- gptcache/config.py Zilliz. Accessed .
- gptcache/adapter/adapter.py Zilliz. Accessed .
- SemanticCacheLookup policy (Apigee) Google Cloud. Accessed .
- Get started with semantic caching policies (Apigee) Google Cloud. Accessed .
- Introduction to semantic cache (Azure Cosmos DB) Microsoft. Accessed .
- Build a read-through semantic cache with Amazon OpenSearch Serverless and Amazon Bedrock Amazon Web Services. Published . Accessed .
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its Countermeasures (NDSS 2026) NDSS Symposium. Accessed .
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching (arXiv 2601.23088v2, ICML 2026) arXiv. Published . Accessed .
- The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems (arXiv 2409.20002v5) arXiv. Published . Accessed .
- OWASP GenAI LLM Top 10 2026 OWASP Gen AI Security Project. Published . Accessed .
- Cache responses to large language model API requests (llm-semantic-cache-store policy) Microsoft. Accessed .
