Field notes
October 6, 2026
Self-Hosting an Open Model vs. Paying for an API: When Each One Is the Cheaper Bill
Self-hosting tends to win when the machines stay busy. A hosted API tends to win when the work comes in bursts.
If you're trying to decide whether to run an open-weight model on your own hardware or keep paying a hosted API, here is the short answer.
Self-hosting tends to win when the machines stay busy. A hosted API tends to win when the work comes in bursts.
The rest of this post is why, and how to figure out which one you are.
The weights are free. The machine isn't.
Open-weight models changed who gets to run a model. You can download one, put it on a GPU you rent or own, and serve it yourself. No per-token meter running.
That's where a lot of the reasoning stops. The weights are free, so the answers must be cheaper.
But the GPU bills by the hour, or you paid for it up front. Either way, it costs about the same whether it's answering requests or sitting idle between them. Then add power, the serving software, a second machine if you need the service to stay up when one fails, and the person who gets the call when it breaks.
A hosted API works the other way around. You pay per token or per request. When nothing is running, you pay close to nothing.
So you aren't really comparing two prices. You're comparing two shapes of bill. One is mostly fixed. The other is mostly pay-as-you-go.
The one number that decides it
The comparison that matters isn't the price per million tokens on two pricing pages. It's this:
What your hardware costs for an hour, with everything included, divided by the tokens it actually produced in that hour.
The important word is "actually." Not the peak speed from someone's benchmark. What your machine did, with your traffic, at your quality and latency bar.
You don't need real prices to see how this works. Say your machine runs at full capacity for an hour. Each token costs some amount. Now say it runs at half capacity. The hour cost the same, and you got half the work, so each token cost twice as much. At one fifth of capacity, five times as much.
The API's price per token doesn't change based on how busy you are. So somewhere there is a level of use where the two lines cross. Above it, your own hardware is cheaper per token. Below it, the API is.
Where that crossing sits depends on what you pay for hardware, which model you run, how fast the answers have to come back, and which API you're comparing against. There isn't one break-even point that holds for everybody. If someone hands you a single number for all cases, treat it as a guess.
Busy machines and bursty work
Picture two teams that send the same number of requests per day.
The first team's requests arrive steadily. Maybe it's a big batch job that runs through the night, or a steady stream of documents to process. Their GPU stays busy most hours.
The second team's requests arrive in clumps. Quiet for most of the day, then a rush when users show up. Or an agent kicks off a long chain of calls, runs hard for a few minutes, and goes silent. This team has to size its hardware for the rush. Then it pays for that hardware through all the quiet.
Same daily count. Very different bills. That's why "requests per day" is a misleading number on its own. It tells you how much work there is, not how it's spread out.
Agent workloads are often the second kind. Lots of turns, hard to predict, uneven. That's the pattern where a metered API is kindest to you, because it charges for the rush and almost nothing for the quiet.
Serverless GPU services, the kind that shut down when idle, soften this problem. They don't erase it. There's usually a window after the last request before the machine stops billing. There's a cold start when it wakes up, and time spent loading the model. It's a different shape of bill, not a free one. Read the fine print on what counts as billable time.
When your own hardware is the right bill
Self-hosting makes sense under specific conditions. It's a fit, not a personality.
High, steady volume. The machines stay busy most of the day, and you can show it with real numbers, not hopes.
A predictable pattern. You can forecast your baseline. "It might take off next Tuesday" is not a baseline.
A smaller model that's already good enough. If a fine-tuned smaller open model handles your task well, you're paying for a modest machine. You aren't trying to rebuild a frontier lab in a closet.
Rules that outrank cost. Sometimes the data has to stay inside your own network. Sometimes you need guaranteed response times, or a model no hosted provider offers. In those cases "allowed" matters more than "cheaper," and that's a fair reason to self-host.
A mixed plan. Run the high-volume, easy work on your own hardware. Send the hard cases to a hosted frontier model. That keeps your machines busy on the bulk without sizing them for every difficult request that might come in.
When the API is the right bill
Low or spiky volume. If the machine would sit dark most of the day, the per-use bill almost always wins.
You don't know your traffic yet. Stay metered until you have real traces. Don't buy hardware for a guess.
The hard cases need a frontier model. Pay the API for that slice. Don't keep idle machines around just in case.
It's early. When a product is new, being able to change things quickly is worth more than owning the stack.
The lines people leave off the spreadsheet
People. Someone has to keep the serving software running, deal with crashes when memory runs out, rerun your quality checks every time the model changes, and carry the pager. Even part of one engineer's time moves the crossover point. A self-hosting plan with no owner isn't a plan. It's a future outage.
The right comparison. Beating the price of an expensive frontier model is much easier than beating a cheap hosted API that serves the same open model you were going to run yourself. "Self-host vs. API" means nothing until you name which API.
Redundancy. If you need two machines so the service stays up when one fails, you've doubled your capacity before counting any headroom. Unless your traffic fills both, your use per machine just dropped by half.
Optimistic speed numbers. Benchmark throughput usually assumes the machine is fed a full, continuous load. Real traffic rarely looks like that.
How to check your own case
You don't need to settle this as a matter of principle. You can measure it.
- Trace a real week. Tokens in and out, how many requests run at once, when the peaks and quiet stretches happen.
- Price three options against that same week. A frontier API, a hosted open-weight API, and your own hardware at the level of use you actually saw.
- Add the human line. Put real engineering time in the self-host column.
- Start mixed if you're unsure. Self-host the easy bulk only once it alone keeps a machine busy.
- Redo the math when something moves. Prices change. Models change. Your traffic changes. Nobody wins this one permanently.
The short version
This isn't really a question of open versus closed. Open weights are a license and an option. They aren't a coupon for free answers.
The question is whether the machine earns its hour. If it's busy, owning it can be the cheaper bill. If it mostly waits, the API was cheaper all along, because it only charged you when something ran.
When I'm not working through this kind of arithmetic, I write The Shadows Series, which starts with Shadows of the Forgotten: johndclay.com/books