Skip to content
Cloud Security DeskSearch
Menu

Technical guideAI systems

Lock down self-hosted model servers before exposing them

Built-in credentials in vLLM, SGLang and Triton cover less than their names suggest, Ollama's local API has none, and vLLM and SGLang leave their internal planes unauthenticated. Here is what stays open and what must sit in front.

Published
Sources checked
Next review
Reading time
14 minutes
Coverage
vLLM · SGLang · Ollama · NVIDIA Triton Inference Server · Kubernetes · nginx
A tall white server tower with a column of small ports down its right side. A blue plate covers the top four ports, and a blue cable rising from below stops against it. Below the plate, ten ports stay open, and six cables, five amber and one red, rise from the bottom of the frame and turn into them.
Conceptual illustration: the built-in key covers a few ports at the top, while many routes below it stay open to whoever can reach them.

An implementation guide for platform and ML infrastructure engineers running open-weight inference, built from vLLM, SGLang, Ollama and Triton documentation, source code and advisories reviewed October 9, 2026. It counts the vLLM 0.31.0 routes outside the API key, charts vLLM advisories, separates two exposure studies and sets out a proxy, network and patching bar.

At a glance

Key findings

  • In vLLM 0.31.0, --api-key checks only paths under /v1, /v2, /inference and /cohere; the security guide lists 21 routes that stay open without any opt-in flag, including /invocations, /pause and /update_weights. [1][2]
  • SGLang's key covers every route except health, readiness and metrics, Ollama's local API has no authentication, and Triton's per-group header secrets are a beta feature. [6][7][12]
  • vLLM's key check could be bypassed through the Host header from 0.3.0 until 0.22.0 (CVE-2026-48746, CVSS 9.1), and 9 of the 11 vLLM advisories published on October 9, 2026 named fixes released between March 20 and September 22. [13][15]
  • PyTorch distributed, KV transfer, ZeroMQ, Ray and gRPC planes have no authentication by design; three critical SGLang flaws had no patch when CERT/CC published them, and the 0.5.21 scheduler still unpickles network messages. [1][17][19][20]
  • The 1,139 and 175,108 exposed-Ollama counts come from one Shodan collection and 293 days of Censys scanning, so they cannot be compared or added. [25][26]

What stays open after the key and the private network

On vLLM 0.31.0, released October 5, 2026, --api-key checks a bearer token only on paths that begin with /v1, /v2, /inference or /cohere. The project's security guide lists 21 other routes that exist without any opt-in flag and never ask for the key, among them /invocations, which runs the same inference functions as /v1, and the control routes /pause, /abort_requests and /update_weights. If --host is not set, the 0.31.0 launcher binds the listening socket to every interface. [1][2][3]

The other common servers draw the line somewhere else. SGLang's --api-key covers every route except health, readiness and metrics, and administrative routes can take a second key. Ollama's local API has no authentication at all. Triton protects named API groups with a shared header secret in a feature NVIDIA labels beta. Behind the HTTP port, neither vLLM nor SGLang authenticates traffic between its own processes and nodes: PyTorch distributed rendezvous, NCCL, KV-cache transfer, ZeroMQ sockets, Ray and the optional gRPC listener accept whoever can reach them. A private network narrows who that is. It does not add a credential. [1][6][7][12][19]

So the built-in key is a second check, not the boundary. The boundary is an authenticating reverse proxy that is the only network path to the HTTP port and forwards an explicit list of routes, a separate network for node-to-node traffic, an egress allowlist, and a patch habit that follows releases rather than advisories. The sections below take six common assumptions in turn, answer each from the projects' own documentation, source code and advisories, and end with a minimum bar.

Authentication defaults by server, from project documentation and source at the versions shown, reviewed October 9, 2026. [1][3][4][5][6][7][8][9][11][12]
Server and versionDefault listenerBuilt-in credentialStill open with it set
vLLM 0.31.0All interfaces if --host is unset, port 8000--api-key bearer token21 listed routes, API docs pages, gRPC port if enabled
SGLang 0.5.21127.0.0.1, port 30000--api-key, and --admin-api-key for admin routesPaths starting /health, /ready or /metrics
Ollama 0.40.2127.0.0.1:11434; container image 0.0.0.0None on the local APIEvery route, and signed-in cloud models
Triton (2.73.0 docs)HTTP, gRPC and metrics all enabledHeader secret per API group, betaAny group not listed, and metrics

The vLLM key guards four path prefixes

The first assumption is that setting --api-key or VLLM_API_KEY puts a vLLM server behind a token. The source shows why it cannot. In 0.31.0 the authentication middleware holds a tuple of four guarded prefixes and verifies the bearer token only when the request path starts with one of them. Any other path, and any OPTIONS request, goes straight to the application. [2]

The security guide turns that rule into lists. At the v0.31.0 tag it names 26 routes that require the key, 10 of which exist only when an extra flag or variable is set, such as the render endpoints behind --enable-scale-out or adapter loading behind --enable-lora with VLLM_ALLOW_RUNTIME_LORA_UPDATING. It names 21 routes that never ask for the key and need no opt-in: six inference routes, nine operational controls that appear whenever the served model supports generation, and six utility routes. Another 11 open routes appear only with --enable-tokenizer-info-endpoint, VLLM_SERVER_DEV_MODE=1 or --profiler-config, and the development set includes /collective_rpc, which the guide calls extremely dangerous. [1]

The open inference routes matter most. /invocations is the SageMaker-compatible entry point and reaches the same functions as /v1, while /pooling, /classify, /score, /rerank and /generative_scoring expose other model tasks without credentials. The open controls let anyone who reaches the port pause generation, abort in-flight work, trigger elastic scaling or call /update_weights. Two more gaps sit outside the lists: FastAPI's schema, Swagger and ReDoc pages are served unless --disable-fastapi-docs is set, and the optional gRPC listener behind --grpc-port has no authentication, authorization or encryption. [1][2][4]

Two notes on the count. It comes from the tagged guide, not the documentation site, which tracks the development branch; on October 9 that branch adds one protected route, /v1/systemone, and leaves the membership of the open lists unchanged. And the vllm serve reference warns that the key covers /v1, /v2 and /inference, while the guide and the middleware also guard /cohere. This article follows the source. [1][2][4]

Figure 01

21 vLLM routes stay open with no opt-in flag

At v0.31.0 the security guide lists 26 routes that require the API key and 32 that never do, 21 of which are present without any opt-in flag. [1]

Stacked horizontal bar chart of vLLM 0.31.0 routes by API key coverage, split by whether the guide lists an opt-in flag: protected by the key 16 without and 10 with an opt-in; open inference 6 and 0; open operational control 9 and 0; open utility 6 and 0; open tokenizer info, dev mode and profiler 0 and 11.

Source. Counted from the API Key Authentication Limitations section of the vLLM security guide at tag v0.31.0 (released October 5, 2026), read October 9, 2026. [1]

Method. Each listed route counted once; /abort_requests appears twice in the operational list and is counted once. A route counts as opt-in when the guide names a flag or environment variable it needs. Operational routes are listed as present when the model supports generation. FastAPI docs pages and the gRPC port are not in the lists and are excluded. [1][4]

Accessible table and figure data
Figure 1 accessible table
Route groupNo opt-in flag listedOpt-in flag or variable listed
Protected by the API key1610
Open: inference60
Open: operational control90
Open: utility60
Open: tokenizer info, dev mode, profiler011
Figure 1 accessible table
Route groupNo opt-in flag listedOpt-in flag or variable listed
Protected by the API key1610
Open: inference60
Open: operational control90
Open: utility60
Open: tokenizer info, dev mode, profiler011

SGLang, Ollama and Triton draw the line elsewhere

The second assumption is that what holds for vLLM holds for the other servers. It does not, and each difference changes what the proxy in front has to do.

SGLang inverts vLLM's model. Its documentation gives --host a default of 127.0.0.1 and --port 30000, and the 0.5.21 authentication code requires --api-key on every route once it is set, except paths beginning /health, /ready or /metrics, which stay open so probes and Prometheus need no secret. Administrative routes, such as weight updates and cache flushes, follow a second rule: with --admin-api-key set, only that key opens them and the ordinary key is refused. The case to avoid is the reverse. With only --admin-api-key set, ordinary routes, inference included, take no key at all. [5][6]

Ollama has no server-side credential to turn on. Its API reference says the local API on port 11434 requires no authentication, and the FAQ says the server binds 127.0.0.1 unless OLLAMA_HOST changes it. That default does not survive the official container: the v0.40.2 Dockerfile sets OLLAMA_HOST=0.0.0.0:11434, and Docker publishes a port given without a host address, as in -p 11434:11434, on every host interface. Publish it as -p 127.0.0.1:11434:11434, or on one private address, when a proxy on the same machine is the only intended caller. [7][8][9][10]

A signed-in Ollama server carries a second cost. After ollama signin, the local API also serves cloud models on the owner's account, so an exposed server lends that account to anyone who finds it, the self-hosted counterpart of running managed models on stolen cloud credentials. The FAQ's local-only switch is OLLAMA_NO_CLOUD=1, or disable_ollama_cloud in ~/.ollama/server.json. [7][8]

Triton expects to sit behind a gateway and says so. HTTP, gRPC and metrics are on by default. --http-restricted-api and --grpc-restricted-protocol can demand a header carrying a shared value for named API groups such as model-repository, model-config, shared-memory, statistics and trace, and every group left off the list stays open. NVIDIA marks the feature beta, which argues for treating it as a second check. --model-control-mode defaults to none; the poll and explicit modes allow repository changes at run time, which the guide warns can lead to arbitrary code execution. [11][12]

The key check itself has failed twice

Third, teams assume that once a key is configured, the check itself is sound. vLLM has published two advisories against the check itself. CVE-2025-59425, published October 7, 2025 and rated high at CVSS 7.5, was a string comparison that took longer the more leading characters of a guess were right, which could let an attacker recover the key faster than by brute force; 0.11.0 fixed it. CVE-2026-48746, whose advisory was published June 2, 2026 and rated critical at 9.1, affected every release from 0.3.0 until 0.22.0. The middleware read a path that Starlette rebuilds with the Host header, ASGI servers do not reject / or ? in that header, and so a request could reach a /v1 route while the check saw a different path. [14][15]

The advisory for CVE-2026-48746 says deployments behind an RFC-conforming web server such as nginx were not affected, which is the practical case for never exposing the uvicorn listener directly. In 0.31.0 the middleware reads the raw ASGI path and compares SHA-256 digests of the presented and configured keys in constant time. [2][15]

The wider advisory feed shows how much else changes around that check. Counting published GitHub security advisories by publication date gives 27 in 2025 and 77 from January 1 to October 9, 2026, including 36 in the third quarter alone. Of the 77 entries from 2026, 54 are rated medium, and many describe denial of service from crafted multimodal or sampling input. The count measures disclosure activity. It does not rank vLLM's security against any other server. [13]

Treat the feed as a lagging indicator. The project's policy says an advisory is published when its fix lands in a release, yet 9 of the 11 advisories published on October 9, 2026 name fixed versions from 0.18.0 to 0.30.0, which shipped between March 20 and September 22. A deployment that followed releases already had those fixes. One that waited for advisories went without them for as long as six and a half months. [13][16]

Figure 02

vLLM advisory publications by quarter and severity

Published advisories rose from 27 in 2025 to 77 in 2026 through October 9, most of them medium severity. The count measures disclosure activity, not relative insecurity. [13]

Stacked horizontal bar chart of vLLM GitHub security advisories by quarter of publication and severity: 2025 Q1 1 critical, 1 high, 1 medium, 1 low; Q2 2, 3, 8 and 1; Q3 0, 2, 0 and 0; Q4 0, 4, 3 and 0; 2026 Q1 1, 4, 4 and 0; Q2 1, 2, 11 and 0; Q3 0, 3, 30 and 3; Q4 to October 9 0, 8, 9 and 1.

Source. Counted from the 104 published security advisories of the vllm-project/vllm GitHub repository, read through the GitHub REST API on October 9, 2026. [13]

Method. Grouped by each advisory's published_at date (UTC) and its GitHub severity; no advisory was withdrawn. 2026 Q4 covers October 1 to 9 only. Publication can trail the fixing release by months, so the quarters do not date the fixes. [13][16]

Accessible table and figure data
Figure 2 accessible table
Quarter publishedCriticalHighMediumLow
2025 Q11111
2025 Q22381
2025 Q30200
2025 Q40430
2026 Q11440
2026 Q212110
2026 Q303303
2026 Q4 to Oct 90891
Figure 2 accessible table
Quarter publishedCriticalHighMediumLow
2025 Q11111
2025 Q22381
2025 Q30200
2025 Q40430
2026 Q11440
2026 Q212110
2026 Q303303
2026 Q4 to Oct 90891

A private network is not an authenticated one

Many deployments rest on a fourth assumption: that a private network makes the unauthenticated parts safe. vLLM's guide states the opposite as a design property: every inter-node channel, meaning PyTorch distributed communication, KV-cache transfer and tensor, pipeline and data-parallel traffic, is insecure by default and must be isolated. PyTorch's own policy says its distributed primitives, TCPStore included, carry no authorization protocol, send messages unencrypted and accept connections from anywhere, so anyone with network access can execute code with the privileges of the user running PyTorch. [1][17]

Bind addresses make it worse than it sounds. With TCP initialization, TCPStore listens on all interfaces even on a single host, and the vLLM guide records that the PyTorch team considers this intended. The guide's defaults put the data-parallel master on port 29500 and KV transfer on 14579. A Ray cluster is one trust domain in which cluster access already means code execution on every node, and with RayExecutorV2 vLLM copies the driver's environment variables, tokens included, to Ray workers unless a denylist file excludes them. [1]

SGLang shows how these planes fail. CVE-2026-3059 concerned a ZeroMQ broker in the multimodal generation module that bound to every interface and deserialized whatever arrived. In the 0.5.21 source that broker listens on 127.0.0.1 but still passes each message to pickle.loads(), so the change narrowed who can reach it, not what it trusts. CERT/CC's note VU#777338, released May 18, 2026, added three critical issues in the same runtime, two unauthenticated code-execution paths and a path-traversal file write, with no patch and no vendor response at publication, and the 0.5.21 scheduler still unpickles what its ROUTER socket receives. CERT/CC describes that socket as bound to 0.0.0.0 by default, while SGLang documents 127.0.0.1 as the --host default, so list the sockets the process actually opened rather than trusting either statement. The mitigation CERT/CC offers is to keep these interfaces off untrusted networks. [5][18][19][20]

It follows that the node network should hold the serving processes and nothing else, not the application pods, not notebooks, and not the proxy. Multi-tenant clusters raise a further question these planes do not answer, which is who may share the cache state they carry.

Put an allowlisting proxy in front

vLLM and Triton both call for a proxy or gateway in front of the server, Ollama's FAQ shows how to put one there, and vLLM is specific about its job: allowlist only the endpoints users need, block everything else including the unauthenticated inference and control routes, and add authentication, rate limiting and logging there. In practice the proxy carries five duties, and the nginx fragment below shows most of them in one place. [1][8][11]

  • Forward only the exact paths and methods the application uses, typically POST /v1/chat/completions and GET /v1/models, and return 404 for everything else, so open routes stay unreachable even when nobody remembers they exist. [1]
  • Authenticate callers with your own identity system and keep the server's key away from clients. The proxy adds Authorization: Bearer with a key only it holds, which turns --api-key into proof that a request came through the proxy. [1][2]
  • Overwrite request identifiers. GHSA-vfp2-c8pq-v6h6, published October 9, 2026, showed that vLLM copied a client's X-Request-Id into the engine request ID and that, with P2pNcclConnector, a crafted value could make a prefill worker connect out and send that request's KV-cache tensors to an attacker; 0.24.0 fixed it. A proxy that sets its own identifier removes the input. [21]
  • Replace the client's Host header. nginx sends its upstream host name by default, and the CVE-2026-48746 advisory says deployments behind an RFC-conforming server such as nginx were not affected. [15][22]
  • Bound body size, rate and concurrency before the engine sees a request. vLLM's own cap on the n parameter defaults to 16384 sequences, far more than most applications need. [1]
Example nginx fragment for a vLLM backend, not a complete configuration. Names, addresses and the key are placeholders; inject the key from a secret store at deploy time and never commit it. Warning: any location added without auth_request is open to every caller.
# Example fragment: nginx built with --with-http_auth_request_module.
# Addresses, host names and the key are placeholders.
upstream vllm_api {
    server 10.0.12.7:8000;            # private address of the vLLM server
}

server {
    listen 443 ssl;
    server_name inference.example.com;
    # ssl_certificate and ssl_certificate_key omitted from this fragment
    client_max_body_size 2m;          # size for your largest legitimate request

    location = /_authz {
        internal;
        proxy_pass https://authz.example.internal/check;
        proxy_ssl_verify on;          # nginx does not verify upstream TLS by default
        proxy_ssl_trusted_certificate /etc/nginx/certs/authz-ca.pem;
        proxy_pass_request_body off;
        proxy_set_header Content-Length "";
        proxy_set_header X-Original-URI $request_uri;
    }

    location = /v1/chat/completions {
        limit_except POST { deny all; }
        auth_request /_authz;
        include /etc/nginx/snippets/vllm-upstream.conf;
    }

    location = /v1/models {
        limit_except GET { deny all; }
        auth_request /_authz;
        include /etc/nginx/snippets/vllm-upstream.conf;
    }

    # Everything else, including /invocations, /pause, /update_weights,
    # /metrics and the API docs pages, never reaches vLLM.
    location / {
        return 404;
    }
}

# /etc/nginx/snippets/vllm-upstream.conf
# Keep every proxy_set_header for a location in one place: a level that
# sets any of them inherits none from the level above.
proxy_pass http://vllm_api;
proxy_http_version 1.1;
proxy_buffering off;                          # pass streamed tokens through
proxy_set_header X-Request-Id $request_id;    # never forward the client's value
proxy_set_header Authorization "Bearer REPLACE_FROM_SECRET_STORE_AT_DEPLOY";
Figure 03

One way in, one admin path and a closed node network

Clients reach the server only through the authenticating proxy, operators use a separate path, nodes talk only to each other, and egress goes only to an approved mirror. [1][11][24]

Conceptual architecture. Clients connect to an authenticating proxy, which forwards allowlisted routes with a backend key to the vLLM API server on a private address. Operators reach control routes through a separate operator listener on an admin network. The API server talks to peer GPU nodes over an isolated network for TCPStore, NCCL and KV transfer, and its only egress is to an approved model mirror.

Source. Conceptual architecture based on the vLLM security guide, NVIDIA's Triton secure deployment guidance and the Kubernetes NetworkPolicy documentation. [1][11][24]

Method. Conceptual. Edges show the permitted connections only; anything not drawn is blocked. A design pattern drawn from the cited guidance, not a tested deployment.

Accessible table and figure data
Figure 3 accessible table
ComponentReachable fromControl
ClientsInternet or corporate networkAuthenticate at the proxy
Authenticating proxyClientsExact routes, backend key, own request ID
OperatorsAdmin networkOperator identity
Operator listenerOperators onlyControl routes such as pause
vLLM API serverProxy and operator listenerPrivate address, key as a second check
Peer GPU nodesEach other onlyIsolated network for TCPStore, NCCL, KV transfer
Model mirrorServer egress onlyApproved weights and code
Figure 3 accessible table
ComponentReachable fromControl
ClientsInternet or corporate networkAuthenticate at the proxy
Authenticating proxyClientsExact routes, backend key, own request ID
OperatorsAdmin networkOperator identity
Operator listenerOperators onlyControl routes such as pause
vLLM API serverProxy and operator listenerPrivate address, key as a second check
Peer GPU nodesEach other onlyIsolated network for TCPStore, NCCL, KV transfer
Model mirrorServer egress onlyApproved weights and code

Where the nginx fragment can go wrong

The fragment above uses auth_request, which sends a subrequest to an authorization service and admits the original request only on a 2xx answer. That module is not built into nginx by default and needs --with-http_auth_request_module, so check the package you deploy before relying on it. [23]

The second trap is header inheritance. proxy_set_header directives are inherited from an outer block only when the current block defines none, so a location that sets a single header of its own silently drops the backend key and the request-ID override set elsewhere. The fragment keeps the complete set in one included file for that reason. Test the result rather than reading it: a request to /invocations should return 404 from the proxy, and a direct request from the proxy host to /v1/models without the backend key should return 401 from vLLM. [2][22]

Separate the planes and narrow egress

Network placement does the work the servers cannot. Give the HTTP port one inbound source, the proxy. Put operator access to control routes such as /pause or SGLang's weight updates on a separate listener or network with its own identity check. Put node-to-node traffic on an isolated segment that contains only serving processes, set VLLM_HOST_IP to an address on it, and block every other inbound port, as the vLLM guide recommends. [1]

On Kubernetes, NetworkPolicy can express most of this, with limits that matter here. A policy does nothing without a network plugin that enforces it, a pod created before the plugin processes a policy can start unprotected, and behavior for hostNetwork pods is undefined; the most common implementation treats their traffic as node traffic and ignores pod selectors. A deployment that puts GPU pods on the host network needs node firewall rules instead. A default-deny egress rule also blocks DNS unless DNS is allowed explicitly, which the example does. [24]

Egress needs an allowlist of its own. Weights and code should come from an approved mirror, not the open internet. vLLM also fetches image, audio and video URLs on a caller's behalf, and its guide recommends --allowed-media-domains together with VLLM_MEDIA_URL_ALLOW_REDIRECTS=0 so that callers cannot point the server at cloud metadata endpoints or internal services. [1]

Example NetworkPolicy fragment. Namespaces, labels and ports are placeholders; it assumes a network plugin that enforces NetworkPolicy and pods that do not use host networking. The peer rule admits every port between vLLM pods because distributed backends choose ports at run time.
# Example: only the proxy reaches the vLLM HTTP port, only vLLM pods
# reach each other, and egress is limited to a model mirror and DNS.
# Namespaces, labels and ports are placeholders.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: vllm-serving
  namespace: inference
spec:
  podSelector:
    matchLabels:
      app: vllm
  policyTypes:
    - Ingress
    - Egress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: edge
          podSelector:
            matchLabels:
              app: inference-proxy
      ports:
        - protocol: TCP
          port: 8000
    - from:
        - podSelector:
            matchLabels:
              app: vllm
  egress:
    - to:
        - podSelector:
            matchLabels:
              app: vllm
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: model-mirror
      ports:
        - protocol: TCP
          port: 443
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns        # match your cluster DNS pods
      ports:
        - protocol: UDP
          port: 53
        - protocol: TCP
          port: 53

Turning off remote code has not always held

Fifth, --trust-remote-code at its default of false is taken to mean that code from a model repository cannot run. Four vLLM advisories since December 2025 describe paths where it did: CVE-2025-66448, a model config class that resolved auto_map entries regardless of the flag, fixed in 0.11.1; CVE-2026-22807, auto_map modules loaded during model resolution without checking it, fixed in 0.14.0; CVE-2026-27893, two model files that hardcoded trust_remote_code=True, fixed in 0.18.0; and CVE-2026-90553, a processor loader that passed the flag to a function that ignores it, fixed in 0.28.0. [4][13]

Release 0.31.0 closed a fifth path that applies when the flag is on. A request could set mm_processor_kwargs.code_revision and choose which revision of the processor code Transformers imported (GHSA-h3rc-6mm3-gc2m). The 0.31.0 guide now has API endpoints reject per-request mm_processor_kwargs unless the server starts with --trust-request-mm-kwargs, an option it reserves for trusted clients. [1][13]

The flag is one control inside a supply-chain decision, not the decision. Review and pin what loads, and limit what a successful load can reach. Runtime weight and adapter changes carry the same risk over the network: vLLM's /update_weights is open, its LoRA routes take the key but are documented as unsafe for untrusted clients, SGLang treats weight updates as administrative, and Triton's poll and explicit modes accept new repository code. Keep cache directories private as well, because vLLM loads cache contents without integrity verification, including formats that can execute code. [1][5][11]

Read exposure counts by their method

The sixth assumption usually arrives as a number: 175,000 exposed Ollama servers, easy to set beside Cisco Talos's earlier figure of about 1,100 as though exposure had grown a hundredfold. The two numbers come from studies that measured different things, and neither can be added to or trended against the other. [25][26]

Cisco Talos's count is a snapshot of hosts that one passive index had already catalogued, checked with a single harmless prompt. The SentinelLABS count accumulates every host Censys saw over 293 days, and its authors report that hosts observed exactly once make up 36 percent of the unique total while contributing under 1 percent of observations. What both support is narrower and more useful: unauthenticated Ollama servers are easy to find at scale, since Cisco's tool found more than 1,000 in its first 10 minutes, and a persistent core of about 23,000 hosts accounts for most of the observations. [25][26]

Monitor your own estate on the same mechanics. Scan your public ranges from outside for the default ports, 8000 for vLLM, 30000 for SGLang and 11434 for Ollama, and for the internal ones, 29500 and 14579; Cisco suggests scheduled Shodan alerts or a port scanner for this. From a host that is not the proxy, send a read-only request such as GET /version to the server's private address and expect no connection at all. Alert on proxy log entries for any path outside the allowlist, since those are probes for the open routes. [1][4][5][8][25]

Two Ollama exposure studies and their methods as each publication describes them. Not comparable and not additive. [25][26]
StudyMethodWindowResult
Cisco Talos, September 1, 2025Shodan index search on port 11434, then one benign prompt per hostOne collection; over 1,000 hosts in the first 10 minutes1,139 reachable without authentication; 214 (about 18.8 percent) serving a model
SentinelLABS and Censys, January 29, 2026Censys scan observations accumulated over time293 days175,108 unique hosts and 7.23 million observations in 130 countries; persistent core of about 23,000

A minimum bar before any traffic

Run these checks in order. Each assumes the one before it passed, and a failure stops the rollout until it is fixed.

  • List every socket the server process has open. The HTTP port should be bound to a private address and every other listener to the node network; on vLLM, set --host explicitly rather than accepting the empty default. [3]
  • Confirm the proxy is the only inbound path to the HTTP port and that it returns 404 for /invocations, /pause, /update_weights, /metrics and the API docs pages. [1]
  • Confirm the proxy authenticates callers, injects the backend key and overwrites X-Request-Id, and that vLLM rejects a direct /v1 request without the key. [2][21]
  • Confirm development mode, profiler and tokenizer-info flags are off, gRPC is off unless a client needs it, and the node network holds only serving processes. [1]
  • Confirm remote code is off unless reviewed, weights come from the approved mirror, and media fetching is limited to listed domains. [1][4]
  • On Ollama, confirm cloud features are disabled or the server is signed out; on SGLang, set both keys if you set either. [6][8]
  • Record the release you run and the newest release, and schedule the upgrade without waiting for its advisories. [16]

Method and provenance

Source-led technical analysis of vLLM, SGLang, Ollama and NVIDIA Triton documentation and source code at named releases, the vLLM GitHub security advisories (counted through the GitHub REST API), CERT/CC records, CVE identifiers checked against NVD, nginx, Docker and Kubernetes documentation, and two published exposure studies. Sources were reviewed on October 9, 2026.

No server, cluster or network was deployed or scanned. Route counts come from the vLLM 0.31.0 security guide and can change with any release; behavior of SGLang, Ollama and Triton is bounded to the cited documentation and source at the versions named. Advisory counts reflect publication dates, not fix dates or relative security.

AI assistance. AI assisted research synthesis, source counting, drafting, diagram planning and visual production, with deterministic editorial checks. No personal deployment experience, independent human review or live test is claimed.

Published under the Cloud Security Desk organizational byline. Read the practitioner guide policy.

References

  1. Security (docs/usage/security.md at tag v0.31.0) vLLM project. Accessed .
  2. AuthenticationMiddleware source (authenticate.py at v0.31.0) vLLM project. Accessed .
  3. Server launcher source (launcher.py at v0.31.0) vLLM project. Accessed .
  4. vllm serve CLI reference vLLM project. Accessed .
  5. Server Arguments SGLang project. Accessed .
  6. HTTP auth utilities (auth.py at v0.5.21) SGLang project. Accessed .
  7. API Reference: Authentication Ollama. Accessed .
  8. FAQ Ollama. Accessed .
  9. Dockerfile at v0.40.2 Ollama. Accessed .
  10. Port publishing and mapping Docker. Accessed .
  11. vllm-project/vllm published security advisories vLLM project on GitHub. Accessed .
  12. GHSA-wr9h-g72x-mwhm: API key authentication vulnerable to timing attack (CVE-2025-59425) vLLM project on GitHub. Published . Accessed .
  13. GHSA-94f4-hr76-p5j6: OpenAI API Auth Bypass (CVE-2026-48746) vLLM project on GitHub. Published . Accessed .
  14. vLLM Security Policy (SECURITY.md) vLLM project. Accessed .
  15. PyTorch Security Policy: Using distributed features PyTorch project. Accessed .
  16. VU#777338: SGLang contains two remote code execution and one path traversal vulnerability CERT Coordination Center. Published . Accessed .
  17. GHSA-vfp2-c8pq-v6h6: SSRF via X-Request-Id header drives P2pNcclConnector outbound connection vLLM project on GitHub. Published . Accessed .
  18. Module ngx_http_proxy_module nginx. Accessed .
  19. Module ngx_http_auth_request_module nginx. Accessed .
  20. Network Policies The Kubernetes Authors. Accessed .
  21. Detecting Exposed LLM Servers: A Shodan Case Study on Ollama Cisco Talos. Published . Accessed .
  22. Silent Brothers: Ollama Hosts Form Anonymous AI Network Beyond Platform Guardrails SentinelLABS. Published . Accessed .