The narrative around open-weight models tends to be all-or-nothing: either they have caught up with closed frontier systems or they never will. Both framings miss where the technology is actually being deployed.
Jobs open weights handle well
- Narrow classification. Routing tickets, tagging content, filtering noise.
- Structured extraction. Pulling fields out of invoices, forms and transcripts with a fixed schema.
- High-volume, low-stakes generation. Where per-call pricing on hosted APIs dominates the budget.
- Confidential data. Work that contractually cannot leave a private network.
Where hosted frontier models still lead
Open-ended reasoning, long multi-step agent runs and tasks requiring broad world knowledge remain the domain of the largest hosted systems. Serving a comparable model yourself also means owning GPU capacity, quantisation trade-offs and evaluation — a real cost that rarely appears in benchmark comparisons.
The decision that actually matters
Run the maths on your own traffic. If a task is narrow, repetitive and high-volume, a smaller self-hosted model frequently delivers the same accuracy at a fraction of the total cost. If the task is open-ended and rare, paying per call is usually cheaper than owning hardware.
Comments (0)
Log in to join the discussion
Log InNo comments yet