“Once we realized how much these endpoints revealed, the next question was: How many of them are exposed to the internet?” — Michael Katchinskiy, Lava
CVE-2026-47483 and Nvidia's patch
In September Nvidia released a fix for a high-severity vulnerability in its DCGM (Data Center GPU Manager) Exporter, tracked as CVE-2026-47483 and given an 8.2 CVSS rating. The company corrected the issue in DCGM Exporter version 4.8.2; operators are advised to upgrade to that version or later. The flaw, as reported by Michael Katchinskiy of datacenter security startup Lava, could allow unauthenticated actors to crash the GPU monitoring service.
Scale of exposure: roughly 2,100 hosts, 12,000 GPU UUIDs, $100M in hardware
Lava scanned the internet four times between March and May and found about 2,100 GPU servers openly exposing DCGM Exporter metrics over HTTP. Those hosts revealed roughly 12,000 unique GPU UUIDs belonging to about 300 organizations. Nearly half of the exposed GPUs—5,274, or 44 percent—were located in the United States. Lava estimated the exposed GPUs represented about $100 million in hardware, and included high-end and consumer devices such as Nvidia Blackwell Ultra B300 GPUs, H200s, H100s, and consumer RTX 5090 and 4090 systems.

Your scanner finds 4,000 vulns. Which 12 matter?
Nubivance is a Rapid7 Registered Partner delivering vulnerability management as a service - scanning, risk-based prioritization, and remediation follow-through across IT and OT.
Fix the backlogWhat DCGM Exporter and Node Exporter reveal
DCGM Exporters read telemetry from GPUs on a host: hardware identifiers, utilization, memory usage, power consumption, and error events. In the exposed deployments Lava found, all of those metrics were transmitted in plaintext over HTTP and did not require authentication. The plaintext exposure makes it possible to map GPU infrastructure, identify potentially vulnerable systems, and monitor workload activity by following changes in utilization and other metrics.
Lava also inspected Prometheus Node Exporter deployments. The team found 12,096 public Node Exporter hosts exposing server models, operating systems, firmware versions, hostnames, storage paths, and networking hardware. That level of detail can reveal how environments are built and configured and can be used to match systems to known vulnerabilities.
/debug/pprof, resource exhaustion, and potential disruption to AI workloads
About 25 percent of the exposed DCGM hosts also exposed Go's built-in profiling endpoint, /debug/pprof, which returns runtime performance data such as CPU and memory usage, memory allocations, goroutine states, and blocking events. Lava's write-up notes that with enough concurrent unauthenticated requests the exporter could run out of memory and crash, cutting off visibility into GPU health and activity. The same CPU and memory pressure could affect AI training or inference workloads running on those hosts.
How operators, cloud providers, and adversaries are implicated
- Operators and cloud customers: Lava recommends that Nvidia DCGM Exporter, Node Exporter and Prometheus services should not be directly reachable from the public internet and that they be restricted to authorized monitoring infrastructure. Operators should upgrade DCGM Exporter to version 4.8.2 or later.
- Cloud and GPU providers named by Lava: the exposed monitoring services affected customer infrastructure across several neocloud and GPU cloud providers, including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean. Lava reported findings to the affected providers; Katchinskiy says those providers worked with customers to address the exposures.
- Potential adversaries: the combination of unauthenticated telemetry, plaintext GPU UUIDs, and exposed profiler endpoints creates reconnaissance value and a plausible avenue for disruption—both by monitoring workload activity and by inducing crashes through resource-exhaustion attacks.
The central paradox exposed by Lava's scans is blunt: organizations are investing tens of millions in GPU hardware while monitoring stacks and profiling endpoints remain directly reachable from the internet. That visibility into telemetry and configuration details both aids reconnaissance and, in the presence of CVE-2026-47483, offers a straightforward path to operational disruption. The immediate, actionable step is straightforward: upgrade to Nvidia DCGM Exporter 4.8.2 or later and isolate monitoring endpoints behind authorized infrastructure. Whether that will be enough to close the broader gap in AI infrastructure security remains a question Lava's findings make hard to ignore.
Source: The Register — High-severity Nvidia bug could crash GPU monitoring on exposed servers



