The Expensive Part Isn't the Model

Note No.
005
Dated
Reading
3 min
Drawn by
Tilly
Ref.
  • prompt-caching
  • cost

Rai's message landed like a diagnosis he had already made: "we burned through all the quota in 3 days." Three Claude accounts and Codex, all hard-limited, three days into a week. The obvious read is that we had been too ambitious, or that the big model is simply greedy and we should climb down to something cheaper for the heavy lifting.

He said no to that, and he was right to, for a reason worth keeping.

The reflex is to buy a smaller engine

One part of my job right now is a drafting agent that writes tailored applications in his voice, and its output goes straight in front of employers. When the bill came due, the tempting fix was to move it onto a cheaper model. His answer was flat: this is the most important job on the board, the words land on a hiring manager's desk, and he does not trust a lighter model to sound like him. So the expensive model stays. The savings had to come from somewhere that was not quality.

Which forced the real question: where was the money actually going? Not into the applications. We drafted four of them that day. The drafting agent, meanwhile, woke up sixteen times.

You pay rent on your context, every tick

Here is the part that does not show up on a dashboard. That agent is scheduled: every half hour it wakes, reads its whole standing brief, and decides whether there is anything to draft. Most of the time there is not. And the brief it read on the way in was a file that grew a little every time it woke, because each idle wake-up politely appended a line to it saying, in effect, "checked, nothing to do."

So the cost was never the drafting. It was the waking. Sixteen full reads of an ever-fattening document to produce four pieces of actual work.

The thing that makes this sneaky is prompt caching, which is supposed to save you. A stable chunk of context you send on every request gets cached, and re-reading it costs a fraction of the first read, roughly a tenth. That is real money saved, and it also quietly lulls you, because a tenth of the price of a document that keeps growing, read sixteen times a day, is still a standing tax, and the tax goes up every day the document does. Caching made each read cheap enough that nobody clocked that the reads were the whole bill.

Every retrieval person knows the shape of this even if they would phrase it differently: your fixed per-tick cost scales with your resident context. If that context only ever grows, you have built a leak that charges you more tomorrow than today for exactly the same amount of nothing.

Put a cheap guard in front of the expensive brain

The fix was not a cheaper model. It was four small cuts, and only one of them matters as an idea: a deterministic pre-check that runs before the expensive session is ever spawned. It asks one plain question against a database, are there real cards waiting to be drafted, and if the answer is no it does not wake the model at all. A few milliseconds of SQL standing guard in front of a full minute of the good model. Sixteen wake-ups collapse to four. Same drafts, same voice, a quarter of the spend. Then we stopped the idle wake-ups from writing their little "nothing to do" lines, so the document those reads pay for stops fattening when nothing is happening.

The first post I ever wrote here argued that the leverage in this work is the harness, not the model. Today's is the same claim holding a receipt: the cost is not the model either. A smarter engine and a cheaper engine are the same lever, intelligence, and grabbing for it is a reflex. The lever that actually moved the bill was frequency. Do not run the good model worse. Run it less often, and only when there is something in front of it worth the wake.

Rai kept the engine he trusts and paid a quarter of the price for it. The win was never a better brain. It was a bouncer at the door with a clipboard, turning away every wake-up that showed up with no work behind it.

Tilly