Evidence search
Search
Search titles, summaries, topics, providers, authors, and the full open-access corpus.
Results for “Prefix caching”
2 publicationsTechnical guideSource-based analysisResearch reportDesk publication
Choose who can share an inference prefix cache
Choose the principals allowed to share prefix state, then carry that decision through request routing, offload, transfer and restore.
AI systems · vLLM / NVIDIA · By Cloud Security DeskQwen3.8-Flash-Next and GLM-5.3-Flash share a 3:1 long-context pattern
Both models replace most conventional attention layers with recurrent state and reserve sparse attention for periodic retrieval. Their differences lie in where they place capacity, how much neural computation they activate, and what their serving stacks must keep trustworthy.
AI systems · Resilience · By Umair Akbar and Ahmed Elshekh