Cost is an architecture decision
Frontier model calls are expensive and slow. How routing, caching and smaller models turn an unpredictable bill into an engineered cost curve.
There is a familiar arc to AI projects. The prototype calls the most capable model available for every request, because that is the fastest way to find out whether the idea works. It does work. The prototype becomes the product. Usage grows, and one month someone in finance asks why the model bill looks the way it does.
The usual reaction is to treat cost as an optimisation problem to be solved later, by negotiating prices or trimming prompts. I think that is the wrong frame. The cost of an AI system is mostly determined by its architecture, by how requests flow, which model handles what, and what gets computed more than once. Cost is an architecture decision, and it is cheapest to make early.
Not every request needs a frontier model
The single largest lever is model routing. In most real workloads, requests vary enormously in difficulty. Some need deep multi-step reasoning. Many are classification, extraction, reformatting or short answers that a much smaller and cheaper model handles just as well. Sending all of them to the largest model is like sending every support ticket to your most senior engineer.
A router sits in front of the models and decides where each request goes. The simplest version uses rules: task type, input length, which product feature made the call. A more adaptive version uses a small classifier to estimate difficulty, or tries a cheap model first and escalates when its output fails validation or its confidence signals are weak.
The escalation pattern deserves attention because it is both cheap and safe. A small model attempts the task. Its output is checked: does it parse, does it pass the schema, does it agree with the retrieved evidence. If it passes, you are done at a fraction of the cost. If not, the request goes to the larger model. You only pay frontier prices for the requests that actually need them.
Routing needs evaluation behind it
Routing without measurement is just hoping. Before trusting a cheaper path you need evidence that quality holds, and that means running the same evaluation set through each route and comparing results, not on average but per category of request, because a router that is fine on easy questions and quietly worse on hard ones looks acceptable on an aggregate score.
The same evaluation set lets you re-check routing decisions when models change. Model prices and capabilities move quickly, and a routing table that was right six months ago may now be leaving money or quality on the table. Treat it as configuration you revisit, not code you write once.
Stop paying for the same work twice
The second lever is caching, and AI systems repeat far more work than people expect. Embeddings for the same document get recomputed on every re-index. The same system prompt and reference material are sent with every request. Popular questions get answered from scratch again and again.
Each of these has a fix. Store embeddings keyed by a hash of the content and only recompute what changed. Use the prompt-caching features that providers now offer for long, stable prefixes, which can make repeated context far cheaper. Cache final answers for frequent questions where freshness allows, and consider semantic caching, matching new questions to previous ones by meaning, for high-traffic, low-variance features, with care about when a cached answer is still valid.
Context is a budget
Token count is the unit of cost, and the easiest place to waste tokens is the context window. It is tempting to send everything that might be relevant and let the model sort it out. That is expensive, slower, and often worse, because irrelevant context distracts models.
Good retrieval is therefore a cost control as much as a quality control. Retrieve fewer, better passages. Re-rank before you send. Summarise long histories instead of replaying them. Give each feature an explicit token budget and treat exceeding it as a bug to investigate, not a fact of life.
Make cost observable
You cannot manage what you cannot see, and most teams cannot answer a basic question: which feature, which customer, which kind of request is driving the bill? Instrument every model call with the feature that made it, the model used, the tokens in and out, the latency and whether it was served from cache. Put it in the same tracing system as everything else.
Once cost is visible per feature and per request type, the conversations change. You can see that one feature with modest usage accounts for a large share of spend because of a bloated prompt. You can see the effect of a routing change the day after you ship it. Cost stops being a monthly surprise and becomes an engineering metric like latency.
Design for the cost curve you want
Put these together and the shape of the bill changes. Instead of cost rising in lockstep with usage, every request paying frontier prices for full context, you get a curve you have designed: cheap paths for common work, expensive paths only where they earn their keep, repeated work served from cache, and a dashboard that tells you when something drifts.
None of this requires heroic engineering. It requires treating model choice, context size and caching as architectural decisions with trade-offs, made deliberately and measured continuously. The teams that do this early can afford to keep improving their product. The teams that do not end up choosing between their margins and their roadmap.