“This led to nearly a million dollars in charges before discovery and containment,” Unit 42 reports — a concrete example of how quickly token jacking can convert an experimental AI workload into a corporate line-item disaster.
What token jacking is and why it matters
Unit 42 defines token jacking as theft of API keys — “tokens” — that legitimate developers and services use to access commercial AI platforms. Unlike a username-and-password session, these tokens allow programs to initiate non‑interactive sessions on behalf of an account. With many AI providers billing by the token (the small chunks of input and output that make up prompts and responses), a stolen key can enable attackers to run unmetered consumption that is tallied and billed on a cycle, often without immediate detection.
Transfer stations and the proxy economy
Unit 42 documents a gray‑market ecosystem of intermediary services it calls “transfer stations.” These operators resell access to frontier AI models at a fraction of retail price by running proxy software such as new‑api or one‑api. The proxies perform obfuscation, rotation and authentication of real credentials, billing, model routing, and prompt normalization. Advertisements for these services appear on Chinese‑language marketplaces such as Taobao and promise “seller‑issued custom credits” purchased anonymously. Unit 42 links to an external researcher, Harshal Singh, for a detailed deep dive into that marketplace.

Audit-ready is a season. It shouldn't be.
Evidence in spreadsheets, controls drifting between audits, frameworks multiplying on flat headcount. Nubivance runs continuous compliance on Rapid7 Cyber GRC - SOC 2, HIPAA, ISO 27001, PCI, CMMC.
End the scrambleHow attackers harvest and weaponize credentials
Unit 42 outlines multiple capture techniques. Attackers can harvest privileged corporate developer accounts via information stealers or phishing and then create new API keys, provision models, remove billing limits, and disable usage alerts. They can mine keys from improperly secured file shares and public code repositories. Increasingly, attackers are using supply‑chain techniques: poisoned, self‑propagating npm packages — including campaigns Unit 42 names as Shai‑Hulud and Miasma — that steal credentials from developer environments and then infect downstream builds. The report warns these npm campaigns could supply transfer stations with credentials “for years.”
Real‑world impact: scale, cost, and limited recourse
Unit 42 describes transfer stations producing “tens of millions of API calls per day,” and that such volume can translate into “hundreds of thousands of dollars” in usage fees. In incident responses, the team observed attackers integrate exposed credentials into transfer stations within minutes. The result can be catastrophic: the nearly‑million‑dollar case cited above is presented as an example of charges that accumulated before discovery and containment. Unit 42 also notes that affected organizations have “very little recourse” to recover funds billed by AI providers, and that the cost can derail budgets or bankrupt smaller businesses. Developers who use transfer stations for apparent savings also face risks — prompts routed to inferior models and sessions mined for sensitive data.
Recommended protections and Palo Alto Networks capabilities
Unit 42 lays out concrete defensive measures: implement spending limits with alerts on sudden usage changes; review privileged accounts that can provision resources; migrate from long‑term access keys to short‑term bearer tokens; use an AI gateway plus machine authentication to bind LLM traffic to verified machine identities; establish network boundaries for compute resources; and tightly govern development environments to prevent malicious packages from entering CI/CD pipelines.
The report highlights Palo Alto Networks products as controls that reduce risk: Prisma AIRS AI Gateway (centralized API key management, visibility into model usage and token spend, budget limits), Idira Agentic Identity Security (agent registry, strong authentication, just‑in‑time secret retrieval), Koi Agentic Endpoint Security (discovery and governance of endpoint software and packages), Cortex XDR and XSIAM plus Cortex Cloud Identity Security (CIEM, ISPM, DAG, ITDR) for CI/CD and cloud identity monitoring, and Advanced URL Filtering to block known malicious domains. Unit 42 also promotes its AI Security Assessment and an Incident Response hotline for urgent matters.
What this means for developers, security teams, and procurement
- Developers and CI/CD owners: monitor package provenance and lock down build environments. Unit 42 singles out malicious npm packages as a high‑amplification vector (Shai‑Hulud, Miasma) and recommends holding new package versions until their reputations are established.
- Security and cloud teams: enforce short‑lived tokens, machine authentication, and spending alerts. The report emphasizes that “limitless scaling by default” and patchwork billing controls create the window attackers exploit.
- Procurement and finance leaders: plan for the financial exposure of unmonitored token use. Unit 42 documents cases where back‑billed usage produced sudden, unrecoverable charges and urges policy controls and proactive budget limits.
Unit 42 frames token jacking as an AI‑centric twist on an old problem: stolen credentials plus unmetered compute equals rapid, high‑value fraud. Its incidents and technical descriptions map a clear attack chain — from poisoned packages and exposed repositories to proxy‑based transfer stations and runaway billing — and its recommendations focus on removing attack surface (short‑lived tokens, machine identity), improving detection (spending alerts, gateway visibility), and hardening developer pipelines.
For organizations concerned they may be compromised, Unit 42 offers Incident Response contact numbers for regions worldwide and urges immediate engagement. Read the full Unit 42 analysis here: https://unit42.paloaltonetworks.com/ai-token-jacking/




