AI cost

Microsoft Is Saving $600M by Switching AI Models. What It Means for You.

This week it was reported that Microsoft could cut as much as $600 million in inference costs by moving its Copilot from OpenAI and Anthropic models to a cheaper one, Moonshot's Kimi K3. Set aside the geopolitics, which are real and which I will come back to. The plain lesson underneath the headline is one that applies to a five-person company just as much as to Microsoft: at any scale, the AI model you run is the single biggest lever on what AI costs you.

The model is the lever, not the feature

The same task, handed to different models, costs wildly different amounts. Recent published figures put the cost of running a standard benchmark task at roughly $2.75 on the priciest frontier model and around $0.33 on a cheap open one. That is not a rounding difference. It is an eight-to-one spread for output that, on many everyday tasks, a normal person could not tell apart.

For a small business automating one process, that spread is the difference between a line item nobody notices and a bill that makes you cancel the project. So before you argue about which AI product has the nicest features, the more useful question is which model is doing the work underneath, and what it charges per outcome.

The trap: cheapest per token is not cheapest per outcome

Here is the part the headline skips. Kimi K3, the model Microsoft is eyeing, is not actually the cheapest option. It runs around $0.94 per benchmark task with full reasoning, more than some Western models, and it has a habit of reasoning more, which burns extra tokens. A model that is cheap per token but chatty, or that fails more often and has to retry, can cost you more per finished job than a pricier model that gets it right the first time.

The number that matters is not the price per token. It is the cost per outcome that actually worked.

This is the whole reason my companion project, ProcessRaven, exists. It models the cost of running an AI agent through a real process, for every major model, and shows the cost per successful outcome rather than the sticker price. The new models in this story, Kimi K3, GPT-5.6, Gemini 3.5, and the rest, are all in it as of this week, so you can compare them on your own process instead of on a press release.

Cheap has a second bill

The reason Microsoft's move is news is not the price. It is that the cheap model is a Chinese open-weight model, and the US administration is weighing an order to restrict exactly those models. In other words, the cheapest option came with a governance question attached.

Your version of that is smaller but the same shape. The model you pick is also a decision about where your data goes, what your auditor will accept, and whether a provider you rely on will still be usable in a year. The cheapest number on the page should inform that decision. It should not make it for you.

The saving that is actually yours

Here is the twist for most small and mid-sized businesses. Your biggest AI saving is usually not the model at all. It is not automating a broken process in the first place. If a process is undocumented, done differently every time, and full of exceptions, pointing any model at it just makes the mess run faster and locks the cost in. Fix the process first, so it is clear and repeatable, and the model becomes a knob you can turn rather than a gamble you are taking.

That is the order I work in: fix the process, then, if it is worth automating, you automate something that actually works, and the model choice is a cost decision you can make with real numbers. If you want to see what an agent would cost to run on one of your processes before you commit, the AI cost estimator gives you a figure to plan with in about a minute.

The short version

Microsoft is right that the model you run is the biggest lever on AI cost, and switching can save a fortune at scale. But copy the lesson, not the move. Compare on cost per successful outcome, not sticker price. Treat the cheapest model as a candidate, not a conclusion, because cheap can carry a governance bill. And remember that for most companies the largest saving is upstream of the model entirely: fix the process before you automate it.

Work with me

Make sure it is worth automating before you spend on it

I fix one process, turn it into a playbook your team follows, and tell you honestly where automation pays and where it does not. Fixed price, about two weeks, fully remote.

← All resources