Blog · AI Startup
Tokenmaxxing: showing off with your bill
Olaf Lemmens, Founder NinA AI Agency · May 30, 2026 · 6 min read

Thursday, Kasteel Woerden. The SPS IT Regie Summit. A full room, and me on stage.
This week I gave one of the most fun keynotes I've done so far. No technical deep dive on agents or architectures. Instead, a story about something I can't ignore anymore: humans and AI. And halfway through, I dropped our new NinA slogan in public for the first time.
Less prompting, more thinking.
Four words. But there's a whole week of news, irritation and research behind them. Let me explain why I was on that stage, and why I believe it's not just a slogan.
TL;DR
For months Silicon Valley judged its people on how many AI tokens they burned. Meta took its internal leaderboard offline, Microsoft is cutting Claude Code subscriptions, Uber burned its entire annual budget in four months. Meanwhile a 19-year-old from Leiden built something that makes AI talk more sparingly, and Claude Opus 4.8 came out — smarter with fewer steps. The common thread: the tech keeps getting more efficient, humans keep getting more wasteful. And that's where the real win is. Less prompting, more thinking.
Tokenmaxxing: showing off with your bill
For those who don't know the term yet. Tokenmaxxing is Silicon Valley's newest status symbol. Not your car, not your title, but how many AI tokens you consume. Tokens are the unit AI models compute in, and suddenly they became a measure of productivity. Whoever burns the most tokens is the most "AI-native".
The epicenter was Meta. Employees there burned more than 60 trillion tokens in thirty days. At standard API prices that's about 900 million dollars. In tokens. In a month. There was even an internal leaderboard, dubbed "Claudeonomics", with titles like "Token Legend". Until it leaked and the whole thing was taken offline after the backlash.

To be fair, there's a defensible side too. Jensen Huang, the head of Nvidia, said he'd be "deeply alarmed" if an engineer earning half a million didn't burn at least a quarter of that on tokens. His point: you want people to develop the habit of using AI everywhere. There's logic to that. But forming a habit and judging your people on consumption are two very different things.
The bill arrives
And then the bill arrived.
Last week Fortune ran with the news that tokenmaxxing is over. Not because of an ideological shift, but because of the simple math at the bottom of the invoice. Meta pulled the leaderboard. Microsoft is cutting Claude Code subscriptions for employees in several product divisions. And Uber admitted it had already burned through its entire 2026 token budget in the first four months.

I want to be sharp here, because the story is more nuanced than "companies are returning to humans". Microsoft is mainly pushing its people toward its own Copilot tool instead of Claude Code. Officially an internal-tool argument, but according to reporting mostly because costs kept rising. The sharper point underneath: the economics of having AI do work turn out to be trickier than the early promises suggested. Using AI is sometimes more expensive than paying a person.
And that's striking. Because it's the companies that shouted loudest that AI would replace people that are now hitting their own bill. Microsoft's own AI boss said earlier this year that AI would take over most office jobs within eighteen months. No company is returning to humans out of principle. They're running into costs. That's something very different, and a much more honest story.
Caveman: a 19-year-old from Leiden
While the giants burned millions, someone did exactly the opposite.
Julius Brussee, 19, first-year Data Science & AI in Leiden, built Caveman. A skill that forces AI models to communicate extremely concisely. No pleasantries, no filler, just symbols, arrows and dense output. The brain stays the same, only the mouth gets shorter. 4,100 stars on GitHub in three days. Installation is literally one line.

What I love: it works best between AI agents. If you build agents or chain multiple AI steps together, every small inefficiency adds up. That's where the biggest win is. Our developers use it for that reason, and it sits in our agent workflows.
But I'm also honest about the numbers, because that's part of how I want NinA to communicate. The repo claims about 75% savings. The Dutch press made that 65%. An independent benchmark on real coding tasks landed at 14 to 21%. Still meaningful, but far from the headline. Anyone who tests it will figure that out. So I'd rather quote the honest number than the pretty one.
Nice detail: Brussee gave a Lunch & Learn at AI House Amsterdam earlier this month. Nineteen years old. Remember the name. And Julius, there's always a spot at NinA AI for talent like you ;)
Opus 4.8: smarter with fewer steps
I'm writing this newsletter together with Claude Opus 4.8, which is exactly 1 day old. And yes, it's more efficient, but not in the way you might think.
It's not "fewer tokens everywhere". It's fewer steps for the same result. Independent analyses see fewer intermediate steps on agentic tasks at better performance. Databricks even reported 61% lower token costs than the previous version in their own agent.
But the improvement that struck me most (remember, I've only tested for 1 day) fits my story perfectly: the model has become more honest. It jumps to conclusions less often and claims it's done less often when the evidence is thin. That's not a speed win. That's quality per token instead of volume. That is, literally, more thinking.

The real question is an energy question
And here's where one of my favorite clients comes in: Stichting Wakker Dier.
Wakker Dier doesn't look at token limits for the pennies. They look at it because of energy consumption. And in doing so, they get something Meta with its leaderboard completely missed.
There's an economic principle that fits perfectly here: the Jevons paradox. The more efficient a technology becomes, the more we use of it, so total consumption actually rises. Despite all efficiency gains, data center energy demand is growing about 15% per year. Not despite the efficiency, but because of the volume.
The numbers make it concrete. An AI query pulls about 2.9 Wh, nearly ten times as much as a regular search. A reasoning AI query uses five to twelve times more energy than a standard query. And an agent that takes multiple steps can use up to a thousand times more tokens than a single query.
Add that to a culture that rewards people for consumption, and you see what happens. A more efficient model plus tokenmaxxing leads to more tokens burned, not less. The savings per query evaporate against the volume.
Less prompting, more thinking
This is why I was on that stage in Woerden.
Efficiency at the model level, Opus 4.8, doesn't solve it. Efficiency at the agent level, Caveman, doesn't solve it. As long as humans keep asking more to make the meter run, volume eats every saving.
The only knob that actually flips the Jevons paradox is human judgment. Fewer, better-targeted prompts. Not pumping harder, but thinking more sharply before you ask. That's why "Less prompting, more thinking" isn't a marketing slogan for me. It's the solution to the tokenmaxxing problem, in four words.
The story isn't "AI is wasteful". That's too easy, and it's not true. The story is this: the tech gets smarter every month, humans slowly get wasteful, and that's exactly where the win is.
A full room in Woerden nodded along. Not because I'm against AI — I build agents all day. But because everyone in that room recognized it deep inside: the value isn't in how much you ask. It's in what you think before you ask.
Until next time,
Olaf Lemmens
Founder NinA AI Agency
P.S. Curious how we approach AI projects from a "less prompting, more thinking" angle? Plan an intro call. Or visit www.nina-ai.nl