Somewhere in a finance meeting right now, someone is staring at an AI bill and realizing the robots have picked up a very human habit: they are remarkably good at spending money when nobody hands them a budget.
It rarely starts as one big strategic decision, but rather the way most expensive habits do. You give everyone the tools, you encourage experimentation, and when in doubt you stuff a little more context into the prompt, because surely more information can’t hurt. For a while that feels like progress: the adoption charts climb, people are genuinely using the tools, and everyone gets to say “AI transformation” in a meeting with a straight face.
Then the bill shows up. At Dell Technologies World this spring, the company described a single developer who burned a billion tokens and a $3,400 cloud bill in one day. Uber told its staff to use AI as much as possible, ranked the usage on internal leaderboards, and then watched its 2026 AI budget evaporate in four months, after which it capped employees at $1,500 a month per coding tool. Neither company has a compute problem. They have an inventory problem, and your AI bill is probably the same story in miniature.
A quick note on method: this piece leans on Toyota’s own description of its production system, on prompt-caching and pricing documentation from Anthropic and OpenAI, on a small eval we ran ourselves, and on a few public research datasets we re-plotted rather than re-ran. Every chart is labeled by what it actually is — measured by us, real published data re-plotted by us, or a transparent model. Links are at the end.
Toyota Already Solved This Problem
Toyota’s Just-in-Time system usually gets summarized as making only what is needed, when it is needed, in the amount needed. That sounds like a manufacturing slogan right up until you look at how most of us use AI, at which point it reads less like a slogan and more like a CFO survival guide.
The idea behind Just-in-Time was never “buy less stuff” for its own sake. It was about getting waste out of the flow of work, because excess inventory ties up cash, hides defects, and slows down feedback. You can have a warehouse stuffed with parts and still fail to build the car, simply because the right part wasn’t in the right place at the right time. That maps onto AI almost uncomfortably well. In a factory, excess inventory is parts gathering dust on a shelf. In AI, excess inventory is tokens sitting in context because nobody made a decision about what the model actually needed. Taiichi Ohno, the engineer who built the system, called overproduction the worst of the wastes, precisely because it quietly manufactures all the others.
Ohno goes on to name three enemies — muda, mura, and muri, meaning waste, unevenness, and overburden. Chances are, your AI bill is quietly losing to all three at once: you pay for tokens that create no value (muda), your spend swings wildly from one run of the same task to the next (mura), and you push the model past the point where more context helps (muri). Most teams only ever notice the first one — the obvious waste — and quietly pay for the other two.
The AI version of inventory is context, which means your AI bill isn’t only a technology problem. It’s an operations problem, and operations problems have operations solutions.
Tokenmaxxing Is Overproduction With Better Branding
The reigning AI reflex is tokenmaxxing: if the model can take more context, give it more; if the meeting can be recorded, record it; if the agent can run longer, let it. The instinct is understandable, because more context feels safer, especially when the downside of leaving something out is a confident, fluent, completely wrong answer. Nobody wants to be the person who trimmed the one paragraph the model needed.
The data is less kind to the instinct than the vibes are. A 2026 study of agentic coding tasks found that runs on the very same task varied by up to 30x in total token usage, and, the part worth sitting with, higher token usage did not reliably buy higher accuracy. Past a point, the extra spend simply stops converting into quality. The problem isn’t that agents are expensive, it’s that they’re variably expensive, which is a much harder thing to budget.
Part of the cure is a distinction most of us skip. “Stop dumping everything into context” is not the same advice as “stop capturing things,” because there are excellent reasons to capture more operational data: meeting transcripts, customer calls, support tickets, and product usage can all become valuable raw material. The mistake is treating the raw-material warehouse and the model’s active workbench as the same place. A data lake is allowed to be messy, historical, and comprehensive. A context window should be curated, current, and task-specific. One is the warehouse, the other is the assembly line. Anyone who has built a personal “second brain” already knows how this fails: you save the article, clip the thread, and tag the PDF, your brain logs the act of saving as the act of understanding, and a year later you own three thousand notes and a vague sense of dread. AI doesn’t fix that, it amplifies it, because a model is very good at retrieving from a pile of unprocessed material and not at all good at telling you the pile was junk.
The Toyota analogy maps cleanly onto the layers:
How the Toyota analogy maps to AI. Each lean idea has an AI equivalent, the overproduction trap it invites, and the just-in-time version that fixes it.
We Ran the Experiment, and the Bill Was the Problem, Not the Brain
To illustrate what we mean, we built a small, deliberately boring test. We hid one specific fact — an access code for a particular site — inside a long document, then surrounded it with four nearly identical decoys (the same kind of code for four other sites), so the model couldn’t just grab the first number it saw; it had to find the right one among look-alikes. We asked a cheap, current model, Claude Haiku 4.5, to retrieve it, ran the test at growing context lengths from 2,000 tokens up to 190,000, and priced every call against live 2026 rates. We were, frankly, trying to make the model rot.
It refused. All the way out to 190,000 tokens, accuracy held at 100%. The decoys didn’t trip it, the length didn’t faze it, and the results came back a uniform, reassuring green, the picture of a model acing an exam. Which is exactly the trap.
Because here is what did change across that same sweep: the cost. Same question, same right answer, every single time, yet the 190,000-token version cost about 78 times more to produce than the 2,000-token version did. You weren’t buying a better answer, you were buying the same answer wrapped in 188,000 tokens of context the model never needed, and you were paying carrying cost on every one of them.
Same answer, 78x the bill. A cheap model (Claude Haiku 4.5) stayed 100% accurate from 2,000 to 190,000 tokens, yet the cost of each correct answer rose 78x across that range. Source: South Shore Analytics.
This is the uncomfortable insight — you do not need the model to be wrong for your spend to be broken. A perfect green checkmark on a retrieval test tells you almost nothing about whether the context was worth paying for, which is why an all-green needle-in-a-haystack chart is a vanity metric and not a victory.
In Chroma’s context rot study, they ran a needle test where the question did not share obvious wording with the answer it was hunting for (a meaning-based match, not a keyword match). On that harder version, GPT-4.1’s accuracy fell off a cliff: from 100% at 5,000 tokens to about 18% by 10,000 and 9% by 50,000, before partially recovering. And this was not a one-model quirk: Chroma ran the same test across eighteen models — GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 among them — and every one degraded as the input grew. GPT-4.1 is simply the model whose raw results they released, so it is the one we can honestly re-plot from data; the curve below is one model’s slice of a pattern that held across all eighteen.
More context, worse answers. On a harder retrieval task, GPT-4.1’s accuracy falls from 100% to under 10% as the input grows - and Chroma found the same decline across all 18 models it tested. Source: Chroma, “Context Rot.”
So the two findings fit together rather than fight. On the easy retrieval that fills most enterprise prompts, the model stays accurate and you simply overpay (Figure 1). On harder reasoning and synthesis, the model also gets less reliable as the context grows (Figure 2). Either way, more context is not free, and an all-green needle test on an easy task tells you nothing about either failure mode.
And the unit you are billed in is sneakier than it looks. A 2026 study found that in roughly 32% of head-to-head comparisons, the model with the lower listed price per token actually cost more per task. The driver is hidden “thinking” tokens, billed at the output rate but never shown, and they swing by as much as 900% across models on the very same question. Their own headline example: Gemini 3 Flash lists about 80% cheaper than GPT-5.4, yet costs about 38% more to actually run.
A model’s sticker price doesn’t predict what it costs. In about a third of head-to-head matchups, the cheaper-listed model is the more expensive one to actually run. Gemini 3 Flash lists near the cheapest ($3.50/M) but ranks second-most expensive in
What Just-in-Time AI Actually Looks Like
The fix isn’t mysterious, it’s mostly discipline, and it has a name now: context engineering, the unglamorous work of deciding what the model sees. The vendors selling you the largest context windows increasingly frame context as a finite resource to be budgeted rather than filled.
It starts with retrieval as the pull system, where the model receives context because the task demands it, not because the company happens to have it lying around. Ask about a contract clause and you retrieve the contract, the relevant policy, and maybe the related customer history, not every legal memo written since 2019 on the theory that one of them might help. Then you cache the stable stuff, because system prompts, tool definitions, policy documents, and reusable examples don’t need to be reprocessed from scratch on every call. Research on long-horizon agentic tasks found that prompt caching cut API costs by as much as 41% to 80% across providers in the best cases, with the honest caveat that caching naively can occasionally make latency worse, so it rewards a little thought rather than a switch you flip and forget.
After that you route the work to the right model, because not every task deserves the most expensive reasoning engine, and as the price-reversal data shows, the most expensive-looking model isn’t reliably the costly one anyway. Classification, extraction, and short summarization are forklift jobs, not Formula 1 jobs. You also schedule for cost when the work can wait: both major providers will run the same job at a 50% discount through their batch lanes, which is just the AI version of shipping freight instead of paying for overnight air. Finally, you tune workflows like operations systems rather than art projects: a 2025 study on automated GenAI workflow tuning reported up to 2.8x better generation quality, up to 10x lower cost, and up to 2.7x lower latency on tasks like RAG question-answering and text-to-SQL. Said differently, structure beats vibes, and it isn’t close.
Put concretely, four levers do most of the work, and each one is just a lean idea in disguise:
The four levers of just-in-time AI spend. Each maps to a lean idea.
Three of those four — caching, retrieval, and routing — stack on the same workload, and that stacking is what the chart below shows: an illustrative bill that starts at naive (everything stuffed into context, no optimization), then adds each layer in turn. Batching sits outside the ladder because it only helps when the work can wait.
The savings compound. Layering caching, retrieval, and model routing onto a naive setup stacks the savings, shown here on an illustrative workload at 2026 prices. Source: South Shore Analytics cost model.
The Counterargument Is Real
The fair pushback is that some work genuinely needs a lot of context. Legal review, diligence, incident response, research synthesis, and codebase-wide debugging can all require a broad view, and starving those workflows in the name of efficiency is just a slower way to get a bad answer. This is also exactly where the rot in Figure 2 applies: pile enough genuinely-needed material into a hard reasoning task and models really do degrade, which is the whole argument for keeping context honest rather than merely large. But Just-in-Time was never zero inventory. Toyota still keeps the minimum stock the line needs to keep flowing, so the point was never austerity, it was synchronization. The AI equivalent is context with a reason: use long context when the task has real ambiguity or genuinely scattered evidence, lean on broader retrieval when recall matters more than cost, and pay for the expensive model when the decision justifies it, as long as that’s a choice you made and not a default you inherited from the demo that worked once.
The market is already showing what unmanaged defaults look like. Fireworks AI reported processing roughly 15 trillion tokens a day in April 2026, up from 10 trillion late last year, while the FinOps Foundation’s 2026 survey found the other half of the story: 98% of organizations now actively manage their AI spend, yet 40% still can’t quantify any value from it, and realized return is averaging about 10%, half the 20% they were targeting. That gap isn’t a reason to panic. It’s a reason to manage the line.
The Takeaway
The companies that win with AI won’t simply be the ones buying the most compute. A few will, because there’s always a place for brute force at the frontier, but for most of us the edge will come from applying boring operational discipline to genuinely expensive intelligence. The move is almost embarrassingly concrete: put a token budget on the same line as every other operating cost, meter spend against accepted outcomes rather than tokens burned, cache what repeats, retrieve what matters, route to the cheapest capable model, batch what can wait, and give your agents a stop signal before experimentation quietly becomes an invoice. Your CFO and your engineering lead should be reading the same dashboard, and in most companies right now, they aren’t.
It’s tempting to capture everything, send everything, summarize everything, and call the resulting bill innovation. The better path is the one Toyota mapped out forty years ago: treat tokens like inventory. Only what is needed, when it is needed, in the amount needed. And hold on to the part that surprised even us, because your model can be getting every single answer right and you can still be overpaying by 78x. The green checkmark was never the finish line. The bill is.
Turns out 1980s Toyota discipline maps pretty nicely onto 2026 AI budgets.
If your AI spend is climbing faster than anyone can explain, that’s the conversation we have all day at South Shore. If you want help optimizing your AI cost structure — turning token spend into a managed line item instead of a mystery — reach out and we’ll take a look.
Thanks for reading!
#AI #DataStrategy #Analytics #FinOps #SouthShoreAnalytics