Most AI bills contain large amounts of avoidable spend. None of the techniques below require a worse product.
Caching
Identical or near-identical requests are far more common than teams expect, especially in support and search workloads. An exact-match cache with a short time-to-live removes a surprising share of calls. Semantic caching extends this to paraphrased questions, at some risk of returning a slightly mismatched answer.
Routing
Send easy requests to a small model and hard ones to a large one. A classifier — which can itself be small — decides the route. In practice most traffic is easy, so the savings are substantial.
Prompt trimming
- Drop examples that never change the output.
- Retrieve fewer, better chunks instead of ten mediocre ones.
- Cap output length explicitly; models ramble without a limit.
- Summarise conversation history instead of resending it verbatim.
Batch where latency is not critical
Providers often offer reduced pricing for asynchronous batch processing. Nightly jobs, backfills and bulk classification rarely need an interactive response time.
Measure before and after
Change one thing at a time and check your evaluation set. Cost reductions that quietly degrade quality are not savings; they are deferred churn.
Comments (0)
Log in to join the discussion
Log InNo comments yet