Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag

Setting Up a Local Language Model on Your Own Machine

Run a capable model offline with no per-token cost, and understand the hardware trade-offs before you start.

Running a model locally removes API costs, keeps data on your machine and works without a connection. The cost is hardware, and knowing the trade-offs in advance saves a wasted weekend.

1. Check your hardware

  • Memory is the constraint. A quantised mid-size model needs roughly 6 to 10 GB of usable GPU memory; larger models need considerably more.
  • Apple silicon uses unified memory, so a 16 GB machine handles mid-size quantised models comfortably.
  • Discrete GPU gives faster generation, but only if the model fits in VRAM. Spilling to system memory cuts speed dramatically.

2. Install a runner

Use a packaged runtime rather than building inference from source. A single installer gives you a model manager, a local API endpoint and a command-line interface, which is everything you need to start.

3. Pull a model

Start with a quantised build in the 7-8 billion parameter range. It is fast enough to be usable on modest hardware and capable enough to be informative about what you actually need.

ollama pull llama3.1:8b
ollama run llama3.1:8b

4. Connect your tools

The runner exposes an HTTP endpoint that mirrors a hosted API shape, so most SDKs work by changing the base URL. Set the base URL to your local port and set the API key to any placeholder value.

5. Measure before scaling up

Record tokens per second and time to first token. If generation runs below roughly ten tokens per second, conversation feels slow regardless of answer quality. Run your own evaluation set rather than trusting a general leaderboard, then decide whether to move up in model size or accept the current one.

When local is the wrong choice

Open-ended reasoning and long agent runs still favour the largest hosted models. Use local inference for narrow, high-volume and private work; use a hosted API for the hard cases.

Comments (0)

Log in to join the discussion

Log In

No comments yet