This website uses cookies

Read our Privacy policy and Terms of use for more information.

Summary

Matt Lenhard kept getting paged for the same problem.

While working on an AI gateway, users were finding ways to abuse access to AI models and run up enormous bills. One Saturday night, while Lenhard was out to dinner with his wife, his team was hit for roughly $60,000. They patched the problem. It disappeared. Then it came back.

As Lenhard started talking to friends at other companies, he realized the problem was much larger. He says some were seeing more than $1 million in token fraud per day. The worst incident he encountered was roughly $10 million over 48 hours.

A few days later, he quit his job.

His investigation eventually led him into Chinese forums, Telegram conversations and an underground economy built around discounted access to models from companies including OpenAI, Anthropic and DeepSeek.

At the top are services often called relays or transfer stations. They sell access to AI inference at discounts that can reach roughly 98% below official pricing. Underneath them is a surprisingly sophisticated supply chain involving account pools, free-credit farming, identity-verification services, payment-card infrastructure, token brokers and other sources of discounted or fraudulently obtained access.

Lenhard didn't just observe the market from the outside. He started contacting people operating inside it.

One broker offered him access to roughly $3 million per month of OpenAI inference at a steep discount. Some of the businesses he encountered looked less like underground hackers and more like conventional software companies, complete with enterprise offerings, invoicing and payment terms.

The economics of AI also create a new kind of security problem.

Lenhard calls one version of it denial of wallet. Traditional software abuse might slow down an application or take a website offline. AI applications are different because individual requests can carry a meaningful cost. An attacker who discovers an expensive AI endpoint can potentially make requests simply to burn a company's money.

Lenhard says he has seen companies suffer seven-figure losses before realizing something was wrong.

That experience led him to build Vectoral, a fraud-detection platform designed for AI applications. Vectoral analyzes signals across the user journey, from registration and browser information to the inference request itself, to estimate whether activity is likely to be abusive.

But Lenhard doesn't think the problem ends with token fraud.

As the major AI labs strengthen identity verification and other controls, he expects attackers to move toward less-protected applications built on top of those models. Longer term, he believes another attack surface will emerge: AI agents.

Instead of an agent committing fraud itself, an attacker could potentially hijack an agent and gain access to its budget, permissions or infrastructure.

The result is a new security problem created by a fundamental change in software economics: when software can spend money every time someone uses it, abuse is no longer just traffic.

It can become a financial attack.

Why This Matters

Token theft is no longer a fringe problem. Stripe says one in six attempted sign-ups at AI services running on its platform comes from a bad actor, while free-trial abuse has more than doubled in six months. Across just eight AI companies, Stripe Radar blocked more than 3.3 million risky sign-ups in a single month.

Stripe sees the problem at the signup and payments layer. Matt Lenhard has been investigating what happens further downstream: relays, brokers, account markets and other infrastructure that can turn abused or fraudulently obtained AI access into deeply discounted inference.

Together, they point to a larger shift in software economics. AI usage has a direct marginal cost and those tokens have resale value. That gives attackers a financial incentive to steal, farm and resell access at scale. If that continues, protecting AI inference and eventually AI agents could become a significant new category of fraud and security infrastructure.

Key Numbers

$10M
Worst token fraud incident Lenhard says he encountered over a 48-hour period.

$1M+
Daily token fraud some people in his network told him they were seeing.

98%
Discount at which Lenhard says some services offer access to frontier AI models.

$3M/month
Amount of OpenAI inference one broker offered Lenhard at roughly a 60% discount.

7 Figures
Losses Lenhard says some larger organizations have suffered before discovering AI abuse.

Further Reading

An Inside Look at the Relay Market Powering Token Resellers and Fraud
Lenhard’s June 28 research piece that prompted much of the conversation and documents the underground token economy.
Read Matt Lenhard’s research

Stripe: AI token theft and the expansion of Radar — This gives you the independent industry-scale data: 1 in 6 sign-ups, abuse doubling, 3.3M blocked attempts.
Read Stripe’s announcement

Hacker News discussion of Lenhard’s research
The article generated substantial technical discussion after being submitted to Hacker News. Simon Willison’s page also links directly to that Hacker News thread.
Hacker News discussion

Simon Willison’s analysis
A useful independent explanation of why the relay market matters for anyone exposing LLM-powered applications publicly.
Simon Willison on the token relay market

Timestamps

00:00 The underground market for stolen AI tokens
01:33 $10M in 48 hours and why Matt quit
02:38 How the stolen AI token market works
08:07 Who buys discounted AI inference?
10:01 How Matt chooses what to build
13:11 Going inside the token broker market
17:22 Model distillation and frontier AI
19:33 “Denial of wallet” attacks
23:53 How Vectoral detects AI abuse
26:06 Where AI fraud goes next
29:40 AI regulation, jobs and moving fast
32:25 What’s next for Vectoral

Go Deeper

Research → (Coming Soon)

Vectoral and the Emerging Market for AI Abuse Prevention

Our research into the underground token economy, Vectoral's thesis and what happens if AI abuse becomes a new category of security infrastructure.

Read the Full Transcript

The complete edited conversation with Matt Lenhard, organized by topic and timestamp.