
An implementation guide for platform and ML infrastructure engineers running open-weight inference, built from vLLM, SGLang, Ollama and Triton documentation, source code and advisories reviewed October 9, 2026. It counts the vLLM 0.31.0 routes outside the API key, charts vLLM advisories, separates two exposure studies and sets out a proxy, network and patching bar.
At a glance
Key findings
- In vLLM 0.31.0,
--api-keychecks only paths under/v1,/v2,/inferenceand/cohere; the security guide lists 21 routes that stay open without any opt-in flag, including/invocations,/pauseand/update_weights. [1][2] - SGLang's key covers every route except health, readiness and metrics, Ollama's local API has no authentication, and Triton's per-group header secrets are a beta feature. [6][7][12]
- vLLM's key check could be bypassed through the
Hostheader from 0.3.0 until 0.22.0 (CVE-2026-48746, CVSS 9.1), and 9 of the 11 vLLM advisories published on October 9, 2026 named fixes released between March 20 and September 22. [13][15] - PyTorch distributed, KV transfer, ZeroMQ, Ray and gRPC planes have no authentication by design; three critical SGLang flaws had no patch when CERT/CC published them, and the 0.5.21 scheduler still unpickles network messages. [1][17][19][20]
- The 1,139 and 175,108 exposed-Ollama counts come from one Shodan collection and 293 days of Censys scanning, so they cannot be compared or added. [25][26]
What stays open after the key and the private network
On vLLM 0.31.0, released October 5, 2026, --api-key checks a bearer token only on paths that begin with /v1, /v2, /inference or /cohere. The project's security guide lists 21 other routes that exist without any opt-in flag and never ask for the key, among them /invocations, which runs the same inference functions as /v1, and the control routes /pause, /abort_requests and /update_weights. If --host is not set, the 0.31.0 launcher binds the listening socket to every interface. [1][2][3]
The other common servers draw the line somewhere else. SGLang's --api-key covers every route except health, readiness and metrics, and administrative routes can take a second key. Ollama's local API has no authentication at all. Triton protects named API groups with a shared header secret in a feature NVIDIA labels beta. Behind the HTTP port, neither vLLM nor SGLang authenticates traffic between its own processes and nodes: PyTorch distributed rendezvous, NCCL, KV-cache transfer, ZeroMQ sockets, Ray and the optional gRPC listener accept whoever can reach them. A private network narrows who that is. It does not add a credential. [1][6][7][12][19]
So the built-in key is a second check, not the boundary. The boundary is an authenticating reverse proxy that is the only network path to the HTTP port and forwards an explicit list of routes, a separate network for node-to-node traffic, an egress allowlist, and a patch habit that follows releases rather than advisories. The sections below take six common assumptions in turn, answer each from the projects' own documentation, source code and advisories, and end with a minimum bar.
| Server and version | Default listener | Built-in credential | Still open with it set |
|---|---|---|---|
| vLLM 0.31.0 | All interfaces if --host is unset, port 8000 | --api-key bearer token | 21 listed routes, API docs pages, gRPC port if enabled |
| SGLang 0.5.21 | 127.0.0.1, port 30000 | --api-key, and --admin-api-key for admin routes | Paths starting /health, /ready or /metrics |
| Ollama 0.40.2 | 127.0.0.1:11434; container image 0.0.0.0 | None on the local API | Every route, and signed-in cloud models |
| Triton (2.73.0 docs) | HTTP, gRPC and metrics all enabled | Header secret per API group, beta | Any group not listed, and metrics |
The vLLM key guards four path prefixes
The first assumption is that setting --api-key or VLLM_API_KEY puts a vLLM server behind a token. The source shows why it cannot. In 0.31.0 the authentication middleware holds a tuple of four guarded prefixes and verifies the bearer token only when the request path starts with one of them. Any other path, and any OPTIONS request, goes straight to the application. [2]
The security guide turns that rule into lists. At the v0.31.0 tag it names 26 routes that require the key, 10 of which exist only when an extra flag or variable is set, such as the render endpoints behind --enable-scale-out or adapter loading behind --enable-lora with VLLM_ALLOW_RUNTIME_LORA_UPDATING. It names 21 routes that never ask for the key and need no opt-in: six inference routes, nine operational controls that appear whenever the served model supports generation, and six utility routes. Another 11 open routes appear only with --enable-tokenizer-info-endpoint, VLLM_SERVER_DEV_MODE=1 or --profiler-config, and the development set includes /collective_rpc, which the guide calls extremely dangerous. [1]
The open inference routes matter most. /invocations is the SageMaker-compatible entry point and reaches the same functions as /v1, while /pooling, /classify, /score, /rerank and /generative_scoring expose other model tasks without credentials. The open controls let anyone who reaches the port pause generation, abort in-flight work, trigger elastic scaling or call /update_weights. Two more gaps sit outside the lists: FastAPI's schema, Swagger and ReDoc pages are served unless --disable-fastapi-docs is set, and the optional gRPC listener behind --grpc-port has no authentication, authorization or encryption. [1][2][4]
Two notes on the count. It comes from the tagged guide, not the documentation site, which tracks the development branch; on October 9 that branch adds one protected route, /v1/systemone, and leaves the membership of the open lists unchanged. And the vllm serve reference warns that the key covers /v1, /v2 and /inference, while the guide and the middleware also guard /cohere. This article follows the source. [1][2][4]
21 vLLM routes stay open with no opt-in flag
At v0.31.0 the security guide lists 26 routes that require the API key and 32 that never do, 21 of which are present without any opt-in flag. [1]

Source. Counted from the API Key Authentication Limitations section of the vLLM security guide at tag v0.31.0 (released October 5, 2026), read October 9, 2026. [1]
Method. Each listed route counted once; /abort_requests appears twice in the operational list and is counted once. A route counts as opt-in when the guide names a flag or environment variable it needs. Operational routes are listed as present when the model supports generation. FastAPI docs pages and the gRPC port are not in the lists and are excluded. [1][4]
Accessible table and figure data
| Route group | No opt-in flag listed | Opt-in flag or variable listed |
|---|---|---|
| Protected by the API key | 16 | 10 |
| Open: inference | 6 | 0 |
| Open: operational control | 9 | 0 |
| Open: utility | 6 | 0 |
| Open: tokenizer info, dev mode, profiler | 0 | 11 |
| Route group | No opt-in flag listed | Opt-in flag or variable listed |
|---|---|---|
| Protected by the API key | 16 | 10 |
| Open: inference | 6 | 0 |
| Open: operational control | 9 | 0 |
| Open: utility | 6 | 0 |
| Open: tokenizer info, dev mode, profiler | 0 | 11 |
SGLang, Ollama and Triton draw the line elsewhere
The second assumption is that what holds for vLLM holds for the other servers. It does not, and each difference changes what the proxy in front has to do.
SGLang inverts vLLM's model. Its documentation gives --host a default of 127.0.0.1 and --port 30000, and the 0.5.21 authentication code requires --api-key on every route once it is set, except paths beginning /health, /ready or /metrics, which stay open so probes and Prometheus need no secret. Administrative routes, such as weight updates and cache flushes, follow a second rule: with --admin-api-key set, only that key opens them and the ordinary key is refused. The case to avoid is the reverse. With only --admin-api-key set, ordinary routes, inference included, take no key at all. [5][6]
Ollama has no server-side credential to turn on. Its API reference says the local API on port 11434 requires no authentication, and the FAQ says the server binds 127.0.0.1 unless OLLAMA_HOST changes it. That default does not survive the official container: the v0.40.2 Dockerfile sets OLLAMA_HOST=0.0.0.0:11434, and Docker publishes a port given without a host address, as in -p 11434:11434, on every host interface. Publish it as -p 127.0.0.1:11434:11434, or on one private address, when a proxy on the same machine is the only intended caller. [7][8][9][10]
A signed-in Ollama server carries a second cost. After ollama signin, the local API also serves cloud models on the owner's account, so an exposed server lends that account to anyone who finds it, the self-hosted counterpart of running managed models on stolen cloud credentials. The FAQ's local-only switch is OLLAMA_NO_CLOUD=1, or disable_ollama_cloud in ~/.ollama/server.json. [7][8]
Triton expects to sit behind a gateway and says so. HTTP, gRPC and metrics are on by default. --http-restricted-api and --grpc-restricted-protocol can demand a header carrying a shared value for named API groups such as model-repository, model-config, shared-memory, statistics and trace, and every group left off the list stays open. NVIDIA marks the feature beta, which argues for treating it as a second check. --model-control-mode defaults to none; the poll and explicit modes allow repository changes at run time, which the guide warns can lead to arbitrary code execution. [11][12]
The key check itself has failed twice
Third, teams assume that once a key is configured, the check itself is sound. vLLM has published two advisories against the check itself. CVE-2025-59425, published October 7, 2025 and rated high at CVSS 7.5, was a string comparison that took longer the more leading characters of a guess were right, which could let an attacker recover the key faster than by brute force; 0.11.0 fixed it. CVE-2026-48746, whose advisory was published June 2, 2026 and rated critical at 9.1, affected every release from 0.3.0 until 0.22.0. The middleware read a path that Starlette rebuilds with the Host header, ASGI servers do not reject / or ? in that header, and so a request could reach a /v1 route while the check saw a different path. [14][15]
The advisory for CVE-2026-48746 says deployments behind an RFC-conforming web server such as nginx were not affected, which is the practical case for never exposing the uvicorn listener directly. In 0.31.0 the middleware reads the raw ASGI path and compares SHA-256 digests of the presented and configured keys in constant time. [2][15]
The wider advisory feed shows how much else changes around that check. Counting published GitHub security advisories by publication date gives 27 in 2025 and 77 from January 1 to October 9, 2026, including 36 in the third quarter alone. Of the 77 entries from 2026, 54 are rated medium, and many describe denial of service from crafted multimodal or sampling input. The count measures disclosure activity. It does not rank vLLM's security against any other server. [13]
Treat the feed as a lagging indicator. The project's policy says an advisory is published when its fix lands in a release, yet 9 of the 11 advisories published on October 9, 2026 name fixed versions from 0.18.0 to 0.30.0, which shipped between March 20 and September 22. A deployment that followed releases already had those fixes. One that waited for advisories went without them for as long as six and a half months. [13][16]
vLLM advisory publications by quarter and severity
Published advisories rose from 27 in 2025 to 77 in 2026 through October 9, most of them medium severity. The count measures disclosure activity, not relative insecurity. [13]

Source. Counted from the 104 published security advisories of the vllm-project/vllm GitHub repository, read through the GitHub REST API on October 9, 2026. [13]
Method. Grouped by each advisory's published_at date (UTC) and its GitHub severity; no advisory was withdrawn. 2026 Q4 covers October 1 to 9 only. Publication can trail the fixing release by months, so the quarters do not date the fixes. [13][16]
Accessible table and figure data
| Quarter published | Critical | High | Medium | Low |
|---|---|---|---|---|
| 2025 Q1 | 1 | 1 | 1 | 1 |
| 2025 Q2 | 2 | 3 | 8 | 1 |
| 2025 Q3 | 0 | 2 | 0 | 0 |
| 2025 Q4 | 0 | 4 | 3 | 0 |
| 2026 Q1 | 1 | 4 | 4 | 0 |
| 2026 Q2 | 1 | 2 | 11 | 0 |
| 2026 Q3 | 0 | 3 | 30 | 3 |
| 2026 Q4 to Oct 9 | 0 | 8 | 9 | 1 |
| Quarter published | Critical | High | Medium | Low |
|---|---|---|---|---|
| 2025 Q1 | 1 | 1 | 1 | 1 |
| 2025 Q2 | 2 | 3 | 8 | 1 |
| 2025 Q3 | 0 | 2 | 0 | 0 |
| 2025 Q4 | 0 | 4 | 3 | 0 |
| 2026 Q1 | 1 | 4 | 4 | 0 |
| 2026 Q2 | 1 | 2 | 11 | 0 |
| 2026 Q3 | 0 | 3 | 30 | 3 |
| 2026 Q4 to Oct 9 | 0 | 8 | 9 | 1 |
A private network is not an authenticated one
Many deployments rest on a fourth assumption: that a private network makes the unauthenticated parts safe. vLLM's guide states the opposite as a design property: every inter-node channel, meaning PyTorch distributed communication, KV-cache transfer and tensor, pipeline and data-parallel traffic, is insecure by default and must be isolated. PyTorch's own policy says its distributed primitives, TCPStore included, carry no authorization protocol, send messages unencrypted and accept connections from anywhere, so anyone with network access can execute code with the privileges of the user running PyTorch. [1][17]
Bind addresses make it worse than it sounds. With TCP initialization, TCPStore listens on all interfaces even on a single host, and the vLLM guide records that the PyTorch team considers this intended. The guide's defaults put the data-parallel master on port 29500 and KV transfer on 14579. A Ray cluster is one trust domain in which cluster access already means code execution on every node, and with RayExecutorV2 vLLM copies the driver's environment variables, tokens included, to Ray workers unless a denylist file excludes them. [1]
SGLang shows how these planes fail. CVE-2026-3059 concerned a ZeroMQ broker in the multimodal generation module that bound to every interface and deserialized whatever arrived. In the 0.5.21 source that broker listens on 127.0.0.1 but still passes each message to pickle.loads(), so the change narrowed who can reach it, not what it trusts. CERT/CC's note VU#777338, released May 18, 2026, added three critical issues in the same runtime, two unauthenticated code-execution paths and a path-traversal file write, with no patch and no vendor response at publication, and the 0.5.21 scheduler still unpickles what its ROUTER socket receives. CERT/CC describes that socket as bound to 0.0.0.0 by default, while SGLang documents 127.0.0.1 as the --host default, so list the sockets the process actually opened rather than trusting either statement. The mitigation CERT/CC offers is to keep these interfaces off untrusted networks. [5][18][19][20]
It follows that the node network should hold the serving processes and nothing else, not the application pods, not notebooks, and not the proxy. Multi-tenant clusters raise a further question these planes do not answer, which is who may share the cache state they carry.
Put an allowlisting proxy in front
vLLM and Triton both call for a proxy or gateway in front of the server, Ollama's FAQ shows how to put one there, and vLLM is specific about its job: allowlist only the endpoints users need, block everything else including the unauthenticated inference and control routes, and add authentication, rate limiting and logging there. In practice the proxy carries five duties, and the nginx fragment below shows most of them in one place. [1][8][11]
- Forward only the exact paths and methods the application uses, typically
POST /v1/chat/completionsandGET /v1/models, and return 404 for everything else, so open routes stay unreachable even when nobody remembers they exist. [1] - Authenticate callers with your own identity system and keep the server's key away from clients. The proxy adds
Authorization: Bearerwith a key only it holds, which turns--api-keyinto proof that a request came through the proxy. [1][2] - Overwrite request identifiers. GHSA-vfp2-c8pq-v6h6, published October 9, 2026, showed that vLLM copied a client's
X-Request-Idinto the engine request ID and that, withP2pNcclConnector, a crafted value could make a prefill worker connect out and send that request's KV-cache tensors to an attacker; 0.24.0 fixed it. A proxy that sets its own identifier removes the input. [21] - Replace the client's
Hostheader. nginx sends its upstream host name by default, and the CVE-2026-48746 advisory says deployments behind an RFC-conforming server such as nginx were not affected. [15][22] - Bound body size, rate and concurrency before the engine sees a request. vLLM's own cap on the
nparameter defaults to 16384 sequences, far more than most applications need. [1]
# Example fragment: nginx built with --with-http_auth_request_module.
# Addresses, host names and the key are placeholders.
upstream vllm_api {
server 10.0.12.7:8000; # private address of the vLLM server
}
server {
listen 443 ssl;
server_name inference.example.com;
# ssl_certificate and ssl_certificate_key omitted from this fragment
client_max_body_size 2m; # size for your largest legitimate request
location = /_authz {
internal;
proxy_pass https://authz.example.internal/check;
proxy_ssl_verify on; # nginx does not verify upstream TLS by default
proxy_ssl_trusted_certificate /etc/nginx/certs/authz-ca.pem;
proxy_pass_request_body off;
proxy_set_header Content-Length "";
proxy_set_header X-Original-URI $request_uri;
}
location = /v1/chat/completions {
limit_except POST { deny all; }
auth_request /_authz;
include /etc/nginx/snippets/vllm-upstream.conf;
}
location = /v1/models {
limit_except GET { deny all; }
auth_request /_authz;
include /etc/nginx/snippets/vllm-upstream.conf;
}
# Everything else, including /invocations, /pause, /update_weights,
# /metrics and the API docs pages, never reaches vLLM.
location / {
return 404;
}
}
# /etc/nginx/snippets/vllm-upstream.conf
# Keep every proxy_set_header for a location in one place: a level that
# sets any of them inherits none from the level above.
proxy_pass http://vllm_api;
proxy_http_version 1.1;
proxy_buffering off; # pass streamed tokens through
proxy_set_header X-Request-Id $request_id; # never forward the client's value
proxy_set_header Authorization "Bearer REPLACE_FROM_SECRET_STORE_AT_DEPLOY";
One way in, one admin path and a closed node network
Clients reach the server only through the authenticating proxy, operators use a separate path, nodes talk only to each other, and egress goes only to an approved mirror. [1][11][24]

Source. Conceptual architecture based on the vLLM security guide, NVIDIA's Triton secure deployment guidance and the Kubernetes NetworkPolicy documentation. [1][11][24]
Method. Conceptual. Edges show the permitted connections only; anything not drawn is blocked. A design pattern drawn from the cited guidance, not a tested deployment.
Accessible table and figure data
| Component | Reachable from | Control |
|---|---|---|
| Clients | Internet or corporate network | Authenticate at the proxy |
| Authenticating proxy | Clients | Exact routes, backend key, own request ID |
| Operators | Admin network | Operator identity |
| Operator listener | Operators only | Control routes such as pause |
| vLLM API server | Proxy and operator listener | Private address, key as a second check |
| Peer GPU nodes | Each other only | Isolated network for TCPStore, NCCL, KV transfer |
| Model mirror | Server egress only | Approved weights and code |
| Component | Reachable from | Control |
|---|---|---|
| Clients | Internet or corporate network | Authenticate at the proxy |
| Authenticating proxy | Clients | Exact routes, backend key, own request ID |
| Operators | Admin network | Operator identity |
| Operator listener | Operators only | Control routes such as pause |
| vLLM API server | Proxy and operator listener | Private address, key as a second check |
| Peer GPU nodes | Each other only | Isolated network for TCPStore, NCCL, KV transfer |
| Model mirror | Server egress only | Approved weights and code |
Where the nginx fragment can go wrong
The fragment above uses auth_request, which sends a subrequest to an authorization service and admits the original request only on a 2xx answer. That module is not built into nginx by default and needs --with-http_auth_request_module, so check the package you deploy before relying on it. [23]
The second trap is header inheritance. proxy_set_header directives are inherited from an outer block only when the current block defines none, so a location that sets a single header of its own silently drops the backend key and the request-ID override set elsewhere. The fragment keeps the complete set in one included file for that reason. Test the result rather than reading it: a request to /invocations should return 404 from the proxy, and a direct request from the proxy host to /v1/models without the backend key should return 401 from vLLM. [2][22]
Separate the planes and narrow egress
Network placement does the work the servers cannot. Give the HTTP port one inbound source, the proxy. Put operator access to control routes such as /pause or SGLang's weight updates on a separate listener or network with its own identity check. Put node-to-node traffic on an isolated segment that contains only serving processes, set VLLM_HOST_IP to an address on it, and block every other inbound port, as the vLLM guide recommends. [1]
On Kubernetes, NetworkPolicy can express most of this, with limits that matter here. A policy does nothing without a network plugin that enforces it, a pod created before the plugin processes a policy can start unprotected, and behavior for hostNetwork pods is undefined; the most common implementation treats their traffic as node traffic and ignores pod selectors. A deployment that puts GPU pods on the host network needs node firewall rules instead. A default-deny egress rule also blocks DNS unless DNS is allowed explicitly, which the example does. [24]
Egress needs an allowlist of its own. Weights and code should come from an approved mirror, not the open internet. vLLM also fetches image, audio and video URLs on a caller's behalf, and its guide recommends --allowed-media-domains together with VLLM_MEDIA_URL_ALLOW_REDIRECTS=0 so that callers cannot point the server at cloud metadata endpoints or internal services. [1]
# Example: only the proxy reaches the vLLM HTTP port, only vLLM pods
# reach each other, and egress is limited to a model mirror and DNS.
# Namespaces, labels and ports are placeholders.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: vllm-serving
namespace: inference
spec:
podSelector:
matchLabels:
app: vllm
policyTypes:
- Ingress
- Egress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: edge
podSelector:
matchLabels:
app: inference-proxy
ports:
- protocol: TCP
port: 8000
- from:
- podSelector:
matchLabels:
app: vllm
egress:
- to:
- podSelector:
matchLabels:
app: vllm
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: model-mirror
ports:
- protocol: TCP
port: 443
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns # match your cluster DNS pods
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
Turning off remote code has not always held
Fifth, --trust-remote-code at its default of false is taken to mean that code from a model repository cannot run. Four vLLM advisories since December 2025 describe paths where it did: CVE-2025-66448, a model config class that resolved auto_map entries regardless of the flag, fixed in 0.11.1; CVE-2026-22807, auto_map modules loaded during model resolution without checking it, fixed in 0.14.0; CVE-2026-27893, two model files that hardcoded trust_remote_code=True, fixed in 0.18.0; and CVE-2026-90553, a processor loader that passed the flag to a function that ignores it, fixed in 0.28.0. [4][13]
Release 0.31.0 closed a fifth path that applies when the flag is on. A request could set mm_processor_kwargs.code_revision and choose which revision of the processor code Transformers imported (GHSA-h3rc-6mm3-gc2m). The 0.31.0 guide now has API endpoints reject per-request mm_processor_kwargs unless the server starts with --trust-request-mm-kwargs, an option it reserves for trusted clients. [1][13]
The flag is one control inside a supply-chain decision, not the decision. Review and pin what loads, and limit what a successful load can reach. Runtime weight and adapter changes carry the same risk over the network: vLLM's /update_weights is open, its LoRA routes take the key but are documented as unsafe for untrusted clients, SGLang treats weight updates as administrative, and Triton's poll and explicit modes accept new repository code. Keep cache directories private as well, because vLLM loads cache contents without integrity verification, including formats that can execute code. [1][5][11]
Read exposure counts by their method
The sixth assumption usually arrives as a number: 175,000 exposed Ollama servers, easy to set beside Cisco Talos's earlier figure of about 1,100 as though exposure had grown a hundredfold. The two numbers come from studies that measured different things, and neither can be added to or trended against the other. [25][26]
Cisco Talos's count is a snapshot of hosts that one passive index had already catalogued, checked with a single harmless prompt. The SentinelLABS count accumulates every host Censys saw over 293 days, and its authors report that hosts observed exactly once make up 36 percent of the unique total while contributing under 1 percent of observations. What both support is narrower and more useful: unauthenticated Ollama servers are easy to find at scale, since Cisco's tool found more than 1,000 in its first 10 minutes, and a persistent core of about 23,000 hosts accounts for most of the observations. [25][26]
Monitor your own estate on the same mechanics. Scan your public ranges from outside for the default ports, 8000 for vLLM, 30000 for SGLang and 11434 for Ollama, and for the internal ones, 29500 and 14579; Cisco suggests scheduled Shodan alerts or a port scanner for this. From a host that is not the proxy, send a read-only request such as GET /version to the server's private address and expect no connection at all. Alert on proxy log entries for any path outside the allowlist, since those are probes for the open routes. [1][4][5][8][25]
| Study | Method | Window | Result |
|---|---|---|---|
| Cisco Talos, September 1, 2025 | Shodan index search on port 11434, then one benign prompt per host | One collection; over 1,000 hosts in the first 10 minutes | 1,139 reachable without authentication; 214 (about 18.8 percent) serving a model |
| SentinelLABS and Censys, January 29, 2026 | Censys scan observations accumulated over time | 293 days | 175,108 unique hosts and 7.23 million observations in 130 countries; persistent core of about 23,000 |
A minimum bar before any traffic
Run these checks in order. Each assumes the one before it passed, and a failure stops the rollout until it is fixed.
- List every socket the server process has open. The HTTP port should be bound to a private address and every other listener to the node network; on vLLM, set
--hostexplicitly rather than accepting the empty default. [3] - Confirm the proxy is the only inbound path to the HTTP port and that it returns 404 for
/invocations,/pause,/update_weights,/metricsand the API docs pages. [1] - Confirm the proxy authenticates callers, injects the backend key and overwrites
X-Request-Id, and that vLLM rejects a direct/v1request without the key. [2][21] - Confirm development mode, profiler and tokenizer-info flags are off, gRPC is off unless a client needs it, and the node network holds only serving processes. [1]
- Confirm remote code is off unless reviewed, weights come from the approved mirror, and media fetching is limited to listed domains. [1][4]
- On Ollama, confirm cloud features are disabled or the server is signed out; on SGLang, set both keys if you set either. [6][8]
- Record the release you run and the newest release, and schedule the upgrade without waiting for its advisories. [16]
Method and provenance
Source-led technical analysis of vLLM, SGLang, Ollama and NVIDIA Triton documentation and source code at named releases, the vLLM GitHub security advisories (counted through the GitHub REST API), CERT/CC records, CVE identifiers checked against NVD, nginx, Docker and Kubernetes documentation, and two published exposure studies. Sources were reviewed on October 9, 2026.
No server, cluster or network was deployed or scanned. Route counts come from the vLLM 0.31.0 security guide and can change with any release; behavior of SGLang, Ollama and Triton is bounded to the cited documentation and source at the versions named. Advisory counts reflect publication dates, not fix dates or relative security.
AI assistance. AI assisted research synthesis, source counting, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.
Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.
References
- Security (docs/usage/security.md at tag v0.31.0) vLLM project. Accessed .
- AuthenticationMiddleware source (authenticate.py at v0.31.0) vLLM project. Accessed .
- Server launcher source (launcher.py at v0.31.0) vLLM project. Accessed .
- vllm serve CLI reference vLLM project. Accessed .
- Server Arguments SGLang project. Accessed .
- HTTP auth utilities (auth.py at v0.5.21) SGLang project. Accessed .
- API Reference: Authentication Ollama. Accessed .
- FAQ Ollama. Accessed .
- Dockerfile at v0.40.2 Ollama. Accessed .
- Port publishing and mapping Docker. Accessed .
- Secure Deployment Considerations (Triton Inference Server) NVIDIA. Accessed .
- HTTP/REST and GRPC Protocols: Limit Endpoint Access (BETA) NVIDIA. Accessed .
- vllm-project/vllm published security advisories vLLM project on GitHub. Accessed .
- GHSA-wr9h-g72x-mwhm: API key authentication vulnerable to timing attack (CVE-2025-59425) vLLM project on GitHub. Published . Accessed .
- GHSA-94f4-hr76-p5j6: OpenAI API Auth Bypass (CVE-2026-48746) vLLM project on GitHub. Published . Accessed .
- vLLM Security Policy (SECURITY.md) vLLM project. Accessed .
- PyTorch Security Policy: Using distributed features PyTorch project. Accessed .
- Multimodal generation scheduler client source (scheduler_client.py at v0.5.21) SGLang project. Accessed .
- VU#777338: SGLang contains two remote code execution and one path traversal vulnerability CERT Coordination Center. Published . Accessed .
- Multimodal generation scheduler source (scheduler.py at v0.5.21) SGLang project. Accessed .
- GHSA-vfp2-c8pq-v6h6: SSRF via X-Request-Id header drives P2pNcclConnector outbound connection vLLM project on GitHub. Published . Accessed .
- Module ngx_http_proxy_module nginx. Accessed .
- Module ngx_http_auth_request_module nginx. Accessed .
- Network Policies The Kubernetes Authors. Accessed .
- Detecting Exposed LLM Servers: A Shodan Case Study on Ollama Cisco Talos. Published . Accessed .
- Silent Brothers: Ollama Hosts Form Anonymous AI Network Beyond Platform Guardrails SentinelLABS. Published . Accessed .