Your per-token rate probably dropped this year. Your bill went up anyway. That's not a billing error, that's how rented inference behaves at scale. Here's the 36-month math, rented API against a model you own, for your team. Move the inputs.
Free to play with. The personalized report is the only thing behind an email.
Renting inference at scale is like leasing a car you drive 200 miles a day. The meter never stops, and it speeds up as you succeed.
| Cost per 1M tokens | Month 12 | Month 36 |
|---|---|---|
| Rented (rate card) | ||
| Owned (marginal) |
Rented cost per token is set by the rate card and never falls for you. Owned marginal cost falls every month you scale. All-in average including hardware and tuning: at month 12, by month 36, still dropping.
Per-token rates on frontier models have fallen as much as 67% since 2024. Enterprise AI invoices went up anyway. Five reasons, none of them on the rate card.
Every price cut gets spent. Cheaper tokens mean more workflows move to AI, which means more tokens. The rate fell 67% and the line item still grew. The unit price was never the variable that mattered.
One 2026 flagship release shipped a tokenizer that produces up to 35% more tokens for the same text. Identical rate card, identical prompts, bigger invoice. There's no negotiation lever for that, because technically the price didn't change.
A chat message was a few thousand tokens. One agent task pushes 400K to 2M tokens through the meter: planning, tool calls, retries, all billed. The shift from chat to agents is why bills jumped 10x while usage felt the same.
Output tokens cost 5x input on every current frontier model, and reasoning models bill their internal thinking as output. The deeper the model thinks, the more you pay, and you don't control how much it thinks.
When a vendor retires a model, every workflow built on it gets re-tested, re-prompted, and re-evaluated on their calendar, not yours. That engineering time never appears on the invoice, but you pay it.
HIPAA, SOC 2, and data-residency rules don't care what a rented model saves you. An owned model runs on your servers or your cloud, inside your access controls. Records and audit trails never leave the perimeter you already had to certify.
You pay per token, forever, on terms you don't control. Rate cuts get eaten by consumption. Tokenizer updates raise the bill without touching the price. Deprecations move your roadmap whether you like it or not.
One capex step when you need capacity, then a flat slope. The model is tuned on your work, runs in your environment, and every improvement accrues to you alone.
The model layer is yours. Fine-tuned on your actual work, your tickets, your documents, your workflows. The weights and the IP live with you, not a vendor. On your specific tasks, a tuned open model routinely holds its own against the frontier, because it isn't competing on everything. It's competing on your thing.
The environment is private. Records, prompts, and audit trails run on your servers or your cloud, inside the access controls you already trust. Nothing about your business transits shared infrastructure.
A rented model gets better for everyone paying the same vendor, including your competitors.
A tuned private model only gets better for you.
Every prompt you send a shared vendor is training data for the market's average. Every prompt you send your own model is compounding an asset you hold.
Curves, break-even, a 12 / 24 / 36-month table, and a written analysis of your team's specific scenario, including whether you should own at all. Delivered as a PDF.
Every default in this calculator is listed and sourced: hardware price, token blend, tuning cost, serving capacity. See the assumptions. Adjust any of them.