Running a model locally removes API costs, keeps data on your machine and works without a connection. The cost is hardware, and knowing the trade-offs in advance saves a wasted weekend.
1. Check your hardware
- Memory is the constraint. A quantised mid-size model needs roughly 6 to 10 GB of usable GPU memory; larger models need considerably more.
- Apple silicon uses unified memory, so a 16 GB machine handles mid-size quantised models comfortably.
- Discrete GPU gives faster generation, but only if the model fits in VRAM. Spilling to system memory cuts speed dramatically.
2. Install a runner
Use a packaged runtime rather than building inference from source. A single installer gives you a model manager, a local API endpoint and a command-line interface, which is everything you need to start.
3. Pull a model
Start with a quantised build in the 7-8 billion parameter range. It is fast enough to be usable on modest hardware and capable enough to be informative about what you actually need.
ollama pull llama3.1:8b
ollama run llama3.1:8b
4. Connect your tools
The runner exposes an HTTP endpoint that mirrors a hosted API shape, so most SDKs work by changing the base URL. Set the base URL to your local port and set the API key to any placeholder value.
5. Measure before scaling up
Record tokens per second and time to first token. If generation runs below roughly ten tokens per second, conversation feels slow regardless of answer quality. Run your own evaluation set rather than trusting a general leaderboard, then decide whether to move up in model size or accept the current one.
When local is the wrong choice
Open-ended reasoning and long agent runs still favour the largest hosted models. Use local inference for narrow, high-volume and private work; use a hosted API for the hard cases.
Comments (0)
Log in to join the discussion
Log InNo comments yet