Ask ten developers which LLM API is cheapest for coding and you will get ten different answers, because each one is measuring something else. One person reads the number on a pricing page. Another looks at last month's invoice. Only a third figure matters: what it costs to get one working, test-passing change out of the model. This article is written in October 2026, so no provider has published a 2027 price list yet. What follows is the method that picks the winner whatever the new numbers turn out to be, plus a clear map of where the cheapest options have been sitting.
The short answer: on sticker price, the cheapest providers are almost always open-weight models served by competing hosts, or labs that price their own models aggressively. On cost per solved task, the winner is rarely the cheapest model on the list. It is the model that passes your tests often enough that retries and review time stay small, with prompt caching switched on. The sections below show why, with a worked example you can rerun using your own numbers.
Why Sticker Price Misleads
A pricing page lists two numbers per model: dollars per million input tokens and dollars per million output tokens. Both are accurate. Both are close to useless for a coding workload when read alone.

Token price is not task price
A model that costs a tenth as much per token but needs three attempts to produce a passing change is not cheap. It is a slow way to spend the same money, plus your own time. Coding is the clearest case for measuring this, because the result can be checked: the tests pass or they do not. That makes cost per passing task the only honest metric, and it is easy to compute once you log tokens and outcomes.
Where the tokens actually go
Agentic coding tools send the whole conversation again on every turn: system prompt, repo files, tool results, earlier edits. A twelve turn session can burn several hundred thousand input tokens to produce a few thousand tokens of code. Two details change the math:
- Output costs more. Output tokens are often priced four to eight times higher than input tokens, so chatty answers and full-file rewrites hurt.
- Thinking counts as output. Reasoning models bill their hidden thinking tokens at the output rate, so a model that looks cheap can cost more than a non-reasoning one on easy edits.
💡 Ask for diffs instead of whole files. A forty line patch costs a fraction of a four hundred line rewrite, and it is easier to review.
Three Types of Providers
Providers fall into three groups, and each group has a different price floor.

Frontier labs, sold direct
OpenAI, Anthropic and Google sell their strongest models directly. You pay the highest sticker price in exchange for the newest releases, the best results on hard multi-file problems, and the most mature caching and batch features. Models such as Claude Sonnet 5, GPT 5.6 Terra and Gemini 3.5 Flash sit in this group. The smaller tiers from each lab, the Flash and mini variants, cost far less than the flagships and handle a surprising share of daily coding work.
Open-weight models on inference hosts
When a model's weights are public, many companies serve it, and they compete on price and speed. GPT OSS 120B, Llama 4 Maverick Instruct and Qwen3 235B A22B Instruct 2507 are examples of families you will find on several hosts. Competition pushes prices toward the raw cost of the GPUs. The catch is variance: hosts differ in quantization, context limits and uptime, so the same model name can behave differently from one host to the next.
Low-price labs and routers
Some labs sell their own models for a fraction of Western flagship prices. DeepSeek V3.1, Kimi K2.6 and Qwen3.7 Plus are the usual names in this group, and through 2025 and 2026 models like them stayed competitive on coding benchmarks while charging far less per token. Router services sit on top of everything, offering one interface and automatic failover across providers. They typically pass list prices through and add a fee for the convenience. Check data handling and server location before sending proprietary code to any of them.
| Provider type | Sticker price | Best at | Watch out for |
|---|
| Frontier lab, direct | Highest | Hard multi-file work, newest models, mature caching | Output price, rate-limit tiers |
| Open-weight host | Low, several hosts compete | Easy switching, predictable pricing | Quantization, context limits, uptime |
| Low-price lab, direct | Usually lowest per token | Strong coding scores for the money | Data location, peak-hour throttling |
| Router service | List price plus a fee | One interface, failover, price comparison | Extra latency, the fee itself |
Four Levers That Cut Your Bill
Before you switch providers, check whether you are using the discounts already on offer. For coding workloads they are worth more than most price differences between vendors.

| Lever | Typical effect | Best for | Trade-off |
|---|
| Prompt caching | Cached input billed at a fraction of the normal rate | Agent loops, big repo context | Cache expires in minutes, prefix must match exactly |
| Batch endpoints | Commonly about half price | Test generation, migrations, evals | Results arrive in hours, not seconds |
| Cheaper model for easy turns | Often 5x to 20x cheaper per token | Inline suggestions, commit messages | Lower solve rate on hard problems |
| Output discipline | Trims the most expensive tokens | Diffs, short explanations | The model must follow the format |
Here is how each one works in practice:
- Prompt caching. Terms differ by provider, but cached input is commonly billed at somewhere between ten and fifty percent of the normal rate. Coding agents are the ideal customer, since they resend the same system prompt and repo context every turn. Some providers cache automatically, while others charge extra to write the cache.
- Batch endpoints. If a job does not need an answer right now, such as generating tests for a hundred files overnight, a batch endpoint commonly halves the price.
- A cheaper model for easy turns. Renaming a variable does not need a flagship. A small model like Claude 4.5 Haiku is built for fast, low-stakes work of this kind.
- Output discipline. Cap output length, ask for diffs, and tell the model to skip the long explanation unless you request it.
💡 Caching only pays when the start of the prompt is identical between calls. Put stable material (system prompt, repo map, style rules) first and the changing task last.
A Worked Cost Example
The prices below are illustrative round numbers, not a quote from any provider. The point is the shape of the calculation, which stays valid when the real 2027 prices arrive.

The setup
- One attempt at a task uses 300,000 input tokens (a twelve turn agent session) and 8,000 output tokens.
- Budget model: $0.30 per million input, $1.20 per million output, solves 50% of tasks on the first attempt.
- Frontier model: $3 per million input, $15 per million output, solves 85% of tasks on the first attempt.
- With caching on, 80% of input tokens are billed at 10% of the normal rate.
- A failed attempt costs the developer 4 minutes of review at $60 per hour, so $4.
Cost per solved task
Divide the cost of one attempt by the solve rate, since a model that succeeds half the time needs two attempts on average.
| Scenario | Cost per attempt | Solve rate | API cost per solved task |
|---|
| Budget, no caching | $0.10 | 50% | $0.20 |
| Frontier, no caching | $1.02 | 85% | $1.20 |
| Budget, caching on | $0.035 | 50% | $0.07 |
| Frontier, caching on | $0.37 | 85% | $0.44 |
On API spend alone the budget model wins by a wide margin, and caching trims every row. If the story ended here, the answer would be simple.
Add the review time
Failed attempts do not just cost tokens. Someone reads the bad output, decides it is wrong and tries again. Using the $4 per failure from the setup:
| Scenario | API cost | Review cost for failures | Total per solved task |
|---|
| Budget, no caching | $0.20 | $4.00 | $4.20 |
| Frontier, no caching | $1.20 | $0.71 | $1.91 |
| Budget, caching on | $0.07 | $4.00 | $4.07 |
| Frontier, caching on | $0.44 | $0.71 | $1.15 |
The ranking flips. The frontier model costs ten times more per attempt and still comes out cheaper per solved task, because its failures are rare. In this example the budget model would need a solve rate of roughly 70% to tie it. That is why "cheapest provider" cannot be answered from a price list alone. To get your own answer:
- Pick 30 real tasks from your repo that have tests.
- Run each candidate model three times in the same agent setup.
- Log input, cached and output tokens, plus pass or fail.
- Build the same two tables with your own prices and review time.
- Repeat every quarter, because prices and models move fast.
Route Models by Task
One model for everything is the expensive default. Match the tier to the difficulty of the work, and the average cost per task drops without hurting quality where it counts.

Easy work: suggestions and tests
Boilerplate, unit test scaffolding, docstrings and commit messages have high solve rates even on small models, so a failure is cheap and rare. Granite 4.1 8B, Claude 4.5 Haiku and GPT 5.6 Luna are the kind of fast, low-cost options to try first.
Medium work: refactors and agent loops
Multi-file refactors and long tool-using sessions are where caching and mid-priced models shine. Gemini 3.5 Flash, DeepSeek V3.1, Kimi K2.6 and Qwen3.7 Plus all deserve a place in your test set. This tier usually carries most of your token volume, so a small price difference compounds here.
Hard work: debugging and design
Race conditions, unfamiliar legacy code and architecture decisions punish weak models: the solve rate drops and the review bill takes over, exactly as in the table above. Spend here. Claude Sonnet 5, GPT 5.6 Sol and Claude Fable 5 are the kind of models to keep for these tasks.

| Task type | Typical share of volume | Model tier | Why |
|---|
| Inline suggestions, tests, docs | High | Small and cheap | High solve rate, cheap failures |
| Refactors, agent loops | Medium | Mid-tier with caching | Most tokens, steady quality |
| Hard debugging, design | Low | Strongest available | Expensive failures |
A simple rule-based router is enough to start. Choose the tier from task labels or file counts, and escalate to a stronger model only after the tests fail twice. Log which tier answered each task. After a month you will see how much volume the cheap tier really absorbed, and you can move the thresholds up or down based on observed failures rather than guesses.
Picking by Team Size
Budget constraints look different for one person than for a team of twenty.
Solo developers

Your time is the scarce resource, and your volume is low. Stay with one mid-tier model for most work, keep a stronger model for the hard problems, and use the allowances that come with your tools before paying per token. Set a monthly spend cap in the provider dashboard on day one.
Small teams

Volume is where the savings are. Agree on a default model per task type, turn on caching everywhere, and move nightly jobs to batch endpoints. A router service can simplify billing and failover, but compare its fee with the savings. Review one monthly report that shows cost per merged pull request, not just total spend. Teams with proprietary code should settle data retention and server location first, because those constraints remove some of the cheapest options before price even comes up. Revisit that monthly report every quarter, since a single new model release can reshuffle the whole table overnight.
Mistakes That Inflate Costs
- No spend cap. A looping agent can run up a large bill overnight. Set hard limits and alerts.
- Sending the whole repo every turn. Give the model a repo map and let it request files.
- No step limit on agent loops. After a set number of failed attempts, stop and escalate to a person or a stronger model.
- Reasoning mode on trivial edits. Thinking tokens bill like output. Switch it off for formatting and renames.
- Switching providers for ten percent. Migration time, prompt rework and quality differences usually erase a saving that small.
Try It on PicassoIA
Before you commit to an API contract, run the same coding prompt on several models and compare the output yourself. PicassoIA gathers dozens of large language models in one place, from small, fast ones to the strongest releases, and you can browse the full catalog at picassoia.com/en/all-models. Paste a real function from your project, ask each model for the same refactor, and note which one needs the fewest follow-up prompts. That small test tells you more than any price list.

The same habit of comparing outputs works for visuals. Blog headers, product mockups and documentation diagrams can be generated in minutes: describe the scene, create several versions with PicassoIA Image, and pick the one that fits. Try creating your own images with Picasso IA today, change one detail in the prompt, and see how much a single sentence moves the result.