A massive global network of 175,000 publicly exposed Ollama AI servers, posing significant remote code execution risks across 130 countries.
An unmanaged layer of AI compute infrastructure operating without the security guardrails and monitoring systems that major platform providers implement by default.
Over a 293-day scanning period, researchers identified 7.23 million observations from unique Ollama hosts spanning 130 countries and 4,032 autonomous system numbers.
The infrastructure analysis revealed a persistent core of approximately 23,000 hosts that generated most of the activity, while a larger layer of transient hosts appeared briefly before disappearing.
Nearly half of the observed hosts are configured with tool-calling capabilities that enable them to execute code, access APIs, and interact with external systems.
This configuration fundamentally changes the threat model beyond simple text generation.
Tool-enabled endpoints can execute privileged operations, and when combined with insufficient authentication and network exposure, this creates what researchers assess as the highest-severity risk in the ecosystem.
The analysis identified 201 hosts running standardized uncensored prompt templates that explicitly remove safety guardrails.
Additionally, 22 percent of hosts support vision capabilities, enabling image understanding and vector creation for indirect prompt-injection attacks via malicious images or documents.
Infrastructure Distribution Patterns
The exposed infrastructure spans both cloud and residential networks globally, challenging traditional assumptions about where AI compute is located.
Fixed-access telecom networks, including consumer ISPs, constitute 56 percent of hosts by count, while hyperscalers account for 32 percent of the infrastructure.
This mixed environment complicates traditional governance approaches that assume centralized control points.
Geographic concentration patterns emerged in the data. In the United States, Virginia accounts for 18 percent of hosts, reflecting cloud infrastructure density in the US-EAST region.
China shows even tighter concentration, with Beijing accounting for 30 percent of hosts and Shanghai and Guangdong together accounting for an additional 21 percent.
Despite decentralized host placement, model adoption shows remarkable concentration.
The same three model families occupy consistent positions across multiple weighting schemes: Llama at number one, Qwen2 at number two, and Gemma2 at number three.
This stability indicates broad, repeated use of shared model lineages rather than fragmented deployment patterns.
Hardware constraints drive convergence toward specific quantization formats. The Q4_K_M format appears on 48 percent of hosts, and 4-bit formats account for 72 percent of all observed quantizations, compared with just 19 percent for 16-bit formats.
This ecosystem-wide convergence creates both portability and fragility, as a vulnerability in how specific quantized models handle tokens could affect a substantial portion of the exposed ecosystem simultaneously.
Security Threat Vectors
This analysis exposed that Ollama’s infrastructure presents multiple distinct threat vectors. Resource hijacking is a primary concern because adversaries can access compute resources without authentication, usage monitoring, or billing controls.
Criminal organizations and state-sponsored actors could leverage these systems for spam campaigns, phishing, or disinformation networks at zero marginal cost.
Prompt injection attacks become increasingly viable with tool-enabled configurations. An attacker can prompt an exposed retrieval-augmented generation instance with benign-sounding requests to extract sensitive information from internal systems.
When vision capabilities are present, indirect prompt injection via malicious images enables sophisticated attacks in which traffic appears to originate from legitimate residential IPs, thereby bypassing standard bot-management defenses.
The residential nature of much of the infrastructure complicates traditional governance approaches and requires new mechanisms that distinguish between managed cloud deployments and distributed edge infrastructure.
Effective incident response relies on clear attribution and centralized control points, but an Ollama instance running on a home network
It may be accessible to adversaries while remaining unreachable by security teams lacking contractual or legal authority.
Emphasize that LLMs deployed at the edge must be treated with the same level of authentication, monitoring, and network controls as other externally accessible infrastructure.
Follow us on Google News , LinkedIn and X to Get More Instant Updates. Set Cyberpress as a Preferred Source in Google.