Same job, 8x the cost: choose AI models by role, not by ranking

The same job can cost 8x more on a flagship model than a lightweight one. How to match model tiers to the work, and how caching and routing protect the budget.

Kai Wu

• Founder, Kaiwu Tech

BusinessPublished Jul 27, 20265 min read

The number first: give the same piece of text work to a flagship model instead of a lightweight one, and the output cost is 8x higher. That is about US$50 per million tokens on one side and about US$6 on the other. For a large share of that work, you cannot tell the two results apart.

This is where most companies actually waste money on AI. Not by using too little of it, but by putting their most expensive staff on their cheapest tasks.

We run AI production pipelines every day, so we learned model selection from our invoices.

The cost is in the job, not the model

Sorting customer-service FAQs, extracting fields from PDFs, filtering logs into reports: give these to a flagship model and the bill is ten times higher, but the accuracy is not ten times better. The difference may sit somewhere after the decimal point, and you still pay ten times as much every month.

The reverse is just as bad. Hand a system refactor that spans many files to a cheap model, and the API fees you save come back tenfold in engineering hours spent cleaning up. Work that is cheap per unit but has to be redone three times is the most expensive work in the company.

So the question was never "is this model expensive?" It is "is this task worth this price?"

The wrong question: "which one is strongest right now?"

This is the question we hear most often. But it is like saying "find out who earns the highest salary on the market and hire them." No company staffs that way. You start by asking what the role does every day, how much capability it needs and what the company can afford to pay.

On top of that, the "strongest" list changes every few months. In the first half of this year alone, the flagship tier turned over once, open-weight models reached trillion-parameter scale, and API costs for the workhorse tier fell by half compared with the previous generation. Tying your company's processes to a single model means betting your staffing strategy on a leaderboard that never stops changing.

The principle worth adopting is counterintuitive: instead of picking the strongest model, draw up a roster that assigns each piece of work its own model.

Three roles: flagship, workhorse, lightweight

Here is how we divide it:

  • Flagship: autonomous tasks that run for days, large codebase refactors, work where a single mistake is costly. The test is whether errors compound. Where they do, getting it right the first time matters far more than saving money.
  • Workhorse: everyday professional work. Organizing documents across systems, producing reports, automating internal processes. This tier usually accounts for 80% of a company's usage, which is exactly why its unit prices deserve the closest scrutiny, item by item.
  • Lightweight: high throughput, low latency. Data preprocessing, bulk classification, filtering for real-time monitoring. Here you are paying for speed and unit price, not depth of reasoning.

There is a fourth role many owners do not think of: the work that stays off the cloud. Medical records, financials, contracts. For regulated data, the question is not which model is smarter. It is whether the data can leave your own server room. Now that frontier-class open-weight models exist, this role has a usable candidate for the first time. Compliance is an architecture decision, not something to patch on later, and model selection is part of the architecture.

Filling in names makes this easier to follow. As of July 2026, the flagship role is held by Claude Fable 5 and GPT-5.6 Sol, the workhorse is GPT-5.6 Terra, the lightweight is GPT-5.6 Luna, and for the off-cloud role, the open-weight Kimi K3 is one of the few usable candidates right now.

But read this list as a snapshot. Next quarter it will certainly look different. The names change. The four roles do not, and neither does the logic you use to assign work. That is exactly why we argue for keeping a routing layer: the definition of a role outlasts the name of any model.

Two ways to protect the budget

The first is caching. Put the fixed system instructions, project background and history into the cache, and the cached portion can cost up to 90% less. For agent workflows that run for long periods, this one item often decides whether a project makes or loses money. Designed well, it can cut operating costs by 60% to 80%.

The second is switching cost. Do not tie your prompts, tool interfaces or entire workflow to one model's quirks. Keep a routing layer in between, so that switching models means changing one setting. You will switch. The only question is whether you do it by choice or because a price increase notice forces you to.

Choosing a model is like hiring: the point was never to find the strongest candidate. It is to have the right-sized person in every role.

This is the same point we made in before you replace employees with AI (in Chinese): AI is not there to replace someone. It is there to reallocate people and budget to where they belong. For how that works in practice, see our case studies.

Three self-checks for owners

  1. What type of work does your company's largest AI bill go to? If you cannot answer, in our experience it is usually going to the cheapest kind of task.
  2. When this work goes wrong once, do you simply rerun it, or does someone have to clean up? If a rerun is enough, give it to a cheap model. Only work that needs a person to clean up justifies a flagship.
  3. If you had to replace your current model today, how many places would you need to change? More than three means you are locked in too tightly, and the next model generation will hurt.

If you are evaluating which AI to adopt, or your bill has stopped making sense, describe how you use it today. We will tell you which parts to downgrade, which are worth upgrading, and how to set up the routing layer in between.

Book a free 30-minute intro call →

Tags

#Small and mid-sized businesses#AI model selection#AI cost control#AI adoption#AI-assisted development

Related posts

More notes on similar problems