Jake Makes AI
The Billing Trick

The Meter Is Always Running

You pay AI by the token, so it has every financial reason to never shut up.

A brass taxi fare meter bolted to a glowing computer monitor, red digits climbing, a blank paper ribbon spilling across the floor

Ask a chatbot a yes-or-no question and count the words it gives back. You asked if a Tuesday launch was a good idea. It returned four paragraphs, a bulleted list of considerations, a "however," a "that said," and a closing offer to help you think through it further. You wanted a yes. You paid for an essay. And somewhere upstream, a meter clicked up a few more cents.

Here is the mechanic almost nobody says out loud. The companies selling AI charge by the token, which is roughly a word-part. Input tokens, output tokens, priced per million. Every product built on top of that API pays the same way. Which means the entire industry sits on top of a pricing model where the vendor makes more money the longer the answer runs. Brevity is a cost center. Rambling is revenue.

Think about what that does to the incentives. If you ran a taxi and got paid by the mile, you would not take the shortcut. You would find a reason to circle the block. The per-token model is a taxi that gets paid by the mile, and it is driving you home the long way every single time, narrating the scenery as it goes.

Brevity is a cost center. Rambling is revenue.

Now stack the reasoning models on top of that. The new hotness is a model that "thinks" before it answers, generating a long internal monologue to work the problem out. Sounds great. Here is the catch you find in the fine print: you pay for those thinking tokens too. The model can burn ten thousand tokens muttering to itself before it writes the two sentences you actually see. You are billed for the muttering. A whole product category got invented whose core feature is that it generates more billable output to answer the same question. Read that again and tell me it is a coincidence.

I am not saying anyone sat in a boardroom and drew up a plan to make the models chatty for the money. It is worse than that, and more honest. Nobody had to. The incentive does the work on its own. When shorter answers cost you less and earn the vendor less, and every metric the vendor optimizes points at usage, nothing in the system is pushing toward the crisp two-word reply. The gradient runs downhill toward more words, and everyone just lets it roll.

You feel it as friction. The preamble that restates your question back to you. The "great question" opener. The disclaimer sandwich, where the actual answer is buried between a paragraph of caveats on each side. The unsolicited summary of what it just said. None of that is helping you. Some of it is safety-team scar tissue, sure. But a lot of it is just filler that happens to be billable, and the thing generating it has no reason on earth to trim it down.

The tell is what happens when the incentive flips. Watch how tight the outputs get the moment a company is paying its own token bill at scale instead of passing it to you. Internal tools that run on the same models suddenly get instructed to answer in one line, no fluff, no preamble, because now the verbosity is coming out of the operator's pocket. Funny how concise the machine can be when the person paying is the one who built it. The capability was always there. The default just points wherever the money is.

So what do you do about it. You stop treating length as effort. A long answer is not a thorough answer, it is an expensive one, and those are different things that got welded together in your head by a thousand bloated replies. You tell the thing to be terse and you mean it. "One sentence. No preamble." You watch your bill by output volume, not just call count, because that is where the leak is. And you get suspicious, deeply suspicious, of any AI product whose pitch is that it thinks longer and writes more. Longer and more is not a feature they built for you. It is a feature they built for the invoice.

The taxi is happy to keep driving. It will tell you all about the neighborhoods you are passing. Just remember who set the meter, and which direction it only ever spins.

§
Post-ready for LinkedIn
Every AI you use is paid by the word. So it has a financial reason to never shut up. Charging by the token is the whole industry's pricing model. Input, output, priced per million. Which means the vendor makes more money the longer the answer runs. Brevity is a cost center. Rambling is revenue. Now look at "reasoning" models. They generate a long internal monologue before answering...and you pay for those thinking tokens too. Ten thousand tokens of muttering to produce two sentences you actually see. A whole product category whose core feature is that it bills more to answer the same question. Nobody had to plan this. The incentive does the work. When shorter answers earn the vendor less, nothing in the system pushes toward the crisp reply. Here's the tell. Watch how tight the outputs get the moment a company pays its own token bill at scale. Internal tools on the same models get told: one line, no preamble. The capability was always there. The default just points wherever the money is. Stop treating length as effort. A long answer isn't a thorough one. It's an expensive one. What's the most useless filler your AI tool adds to every single reply?
← All essays