The rise of good enough inference

For most of the use cases with AI, we don't need the most cutting-edge model.

For most of the use cases with AI, we don't need the most cutting-edge model. What you can do with a good enough model is fantastic, and you can see it in how people use AI. A lot of people are incredibly productive without using the most expensive model from a frontier provider.

I do most of my work in Opus. The frontier model is Fable. I don't need that. Most of my coding work can go over to self-run models, because the cost of running them is so much less.

That doesn't apply to everything you do. If you're doing something one time, like creating code, you want a better model. If you're doing categorization or sorting, you want the most effective model, which may not be the most expensive.

Effective is not the same as expensive

Like everything in business, "most effective" depends on where you're using it. Most of the time it means the most cost-effective correct answer. That comes down to three things: your error rate, how much it costs to get the answer, and how important the question is. Those three together define the right model for the job.

Where we downgraded

Take categorizing interviews. When you're processing text and categorizing interviews, you don't need a deep-thinking model. You need a model that can accurately interpret a long block of text. So we can downgrade: a lower-cost model, an open source model, or a self-hosted one.

We recently moved a lot of our interview transcription to a self-hosted model, so there's no inference cost. It just happens. The accuracy is great for what we need, delivery is fast, and we get it practically free.

What it cost us

It took a lot of testing to decide that was the way to go. One of the great things is that a lot of testing can be pretty cheap to do, and we didn't have to give up much.

What we did take on is an operational and administrative burden we didn't have before. Before, it was just an API call. Now we have to make sure the server is up and running. That's OK, because it runs part-time on a machine we already use. It's not a big deal, and given the expense, the effectiveness is wonderful.

Where the frontier earns its price

I go to a more expensive model when I need the thinking and the analysis. When you're designing a process that will run hundreds of thousands or millions of times, the efficiency and effectiveness of that process is the top priority.

Using a really great model to build the code, the harness, the decision tree can be well worth the money, because you leverage it across so many transactions. Especially if you use that intelligence to drive down the cost of each transaction through how you design it.

Why now

AI changes every day. What we're seeing now is the rise of the good enough model. Models are becoming specialized and easier to tailor to a need, and when a model is tailored to what you need, a lower-cost model can deliver the same results.

You're seeing it in other companies. In legal, Latham & Watkins just brought in its own GPUs to train its own models rather than rely on the frontier providers. I think we'll see that in a number of cases, because the benefit isn't just cost. You keep your secret sauce. You're not sharing your data with somebody who may one day want to compete with you.

It goes back to Satya Nadella's public memo on how important it is for businesses to keep their information their own. In big business, those two trends are coming together.

You don't need a data center

Building a data center is out of scope for most companies. But a lot of companies are purely inference providers, not training providers. They're not as interested in harvesting your data. They want to sell you inference. So if building your own data center or running your own server doesn't fit your company, whether that's budget, capability or strategy, find those providers.

Where good enough goes wrong

In a number of places. One, you misjudge it, and good enough really isn't good enough. The other big one is that whenever you automate a process, you can automate a bad process. At some point this all turns into good change management and choosing the right tools.

In my opinion, it comes back to the human side of AI: engaging the people in the process in the decisions. That way you understand the scope of each decision, the data that goes into it, and you can do a better job of matching that decision and that data to the right level of model.

Start by exploring

The most important thing any company can do is explore. Companies should be using AI to rebuild their processes, and when they do, they need to look at what those processes actually need. You don't need agentic code for everything. You can use AI to build great deterministic code, and operate inside that code with cost-effective models.

Success here looks like it has in every AI transition so far. Look at what's coming. Don't tie yourself to the expensive cutting edge unless you really need it. And take the time to understand your processes, what they really need, and what you can do to accelerate them.


References