Production LLM Optimization: Cutting Cost and Latency by 70%
How model cascading, prompt caching, and structured outputs cut production LLM cost and latency by up to 70% — without sacrificing quality.

The most common mistake I see in production AI systems is brutally simple: every single request goes to the biggest, most expensive model available. A customer asks for opening hours? GPT-4-class model. A webhook payload needs classification? Same model. A complex multi-step reasoning task? Same model again. The result is a monthly bill that scales linearly with traffic and a latency profile that frustrates users on every interaction — even the trivial ones.
In one client audit, 82% of requests were routine tasks that a small model could handle in under 300ms. They were all being served by a frontier model at 2-4 seconds per call and 15x the cost per token. The fix wasn't a better model. It was better routing.
1. Model Cascading & Routing
Not every request deserves your strongest model. The core idea of model cascading is to route each request to the cheapest model that can handle it reliably. Small language models (SLMs) — think 3B to 8B parameter class — are excellent at classification, extraction, summarization of short texts, intent detection, and template-based replies. They respond in a fraction of the time and cost a fraction of the price.
- →Tier 1 — SLM (local or cheap API): intent classification, FAQ matching, data extraction, language detection
- →Tier 2 — Mid-size model: multi-turn conversation, structured summaries, tool selection
- →Tier 3 — Frontier model: complex reasoning, ambiguous requests, anything Tier 1-2 flags as low-confidence
The routing layer itself can be a tiny classifier or even a set of heuristics: message length, presence of specific intents, confidence scores from the SLM. If the small model's confidence drops below a threshold, escalate. In practice, 60-80% of production traffic never needs to leave Tier 1.
"Using a frontier model for every request is like hiring a surgeon to apply band-aids."
2. Prompt Caching & Semantic Caching
Most production prompts have a large static prefix: system instructions, business data, few-shot examples, tool definitions. If you're sending 4,000 tokens of context with every request, you're paying to re-read those 4,000 tokens every single time. Prompt caching changes that: providers now let you cache the stable prefix and pay a fraction of the input cost on cache hits — typically an 80-90% reduction on cached input tokens.
Semantic caching goes one step further. When a user asks a question that's semantically identical to one you've already answered — "what are your hours?" vs "when do you open?" — you can serve the cached answer without calling the model at all. An embedding lookup against a vector store costs milliseconds and fractions of a cent. For FAQ-heavy workloads like customer support, semantic cache hit rates of 30-50% are realistic.
3. Output Tokens & Structured Outputs
Output tokens are the hidden budget killer. They're typically 3-5x more expensive than input tokens, and they generate sequentially — every extra token is extra latency the user feels. A model that answers in 400 tokens when 80 would do isn't being thorough; it's being slow and expensive.
Structured outputs (JSON schemas) solve this elegantly. When you constrain the model to a schema, it stops rambling. Instead of a paragraph of prose you have to parse, you get exactly the fields you asked for — no filler, no pleasantries, no markdown formatting you'll strip anyway. In my benchmarks, switching extraction tasks from free-text to schema-constrained output cut generation time roughly in half and made downstream parsing deterministic.
- →Set explicit max_tokens limits per task type — don't let the model decide
- →Use JSON schemas for any output consumed by code, not humans
- →Ask for terse output in the system prompt: "Answer in under 50 words" works surprisingly well
- →Stream responses so users see output as it generates — perceived latency drops even when total time doesn't
The Practical Result
Stack these three techniques and the numbers compound. Routing 70% of traffic to an SLM cuts that traffic's cost by ~90%. Prompt caching cuts the remaining input cost by 80%+. Structured outputs halve output token spend. Together, a 60-70% reduction in total LLM spend is a conservative estimate — and p95 latency drops in parallel because small models and cache hits are simply faster.
The best part: none of this requires changing your product or degrading the user experience. The user still gets accurate answers. They just get them faster, and you stop paying frontier-model prices for band-aid tasks.
"Optimization isn't about using worse models. It's about using the right model for each request — and not paying for the same tokens twice."