Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Strata Runs a 125-Billion-Parameter Qwen Model on a 12 GB Gaming GPU at up to 93 Tokens per Second

Strata Runs a 125-Billion-Parameter Qwen Model on a 12 GB Gaming GPU at up to 93 Tokens per Second

Strata, a new MIT-licensed inference engine built specifically for Qwen3.8-Flash-Next, runs the 125-billion-parameter MoE model on a single 12-24 GB NVIDIA GPU plus 64 GB of RAM. Developer-published tests show up to 93 tokens per second on an RTX 5070, and an independent r/LocalLLaMA test measured 51 tokens/s versus 23 on stock llama.cpp. Cold experts are computed on the CPU while a 29 GB n-gram table streams from SSD.

A new open-source project is making an unlikely claim: that a 125-billion-parameter frontier-class model can run on the gaming PC under your desk, at speeds faster than most people read. Strata, released on GitHub by developer Niko1221 under the MIT license, is a purpose-built inference engine for Qwen3.8-Flash-Next, Alibaba's 125-billion-parameter mixture-of-experts model. A one-click installer for Windows and Linux handles the entire stack — driver checks, model download of roughly 70 GB, quantization choice and server startup — and exposes both OpenAI-compatible and Anthropic-compatible APIs on localhost, so existing apps and coding agents can use it as a drop-in local provider.

The headline numbers come from the project's own benchmarks, published on the repository and therefore not independently audited. On an RTX 5070 with 12 GB of VRAM, a Ryzen 5 7600 and 64 GB of RAM, the most compressed Q2_0 build writes short-chat answers at 93 tokens per second and sustains 74 tokens per second at 128K context, while reading prompts at up to 2,170 tokens per second. The higher-quality IQ3_S build manages 53 tokens per second in short chat. For cards with more memory the README estimates roughly 100-140 tokens per second on an RTX 3090, a figure the project itself labels as an estimate. The claims have drawn attention well beyond the usual local-LLM circles: tombkeeper, a well-known Chinese security researcher, posted about Strata on Weibo on September 28, citing a test in which an RTX 3090 ran the IQ3_S quantization at 67 tokens per second on short-question workloads where stock llama.cpp managed only 20 on the same hardware — a figure he attributed to community testing.

Crucially, the developer-published numbers are not the only evidence. A user on r/LocalLLaMA ran an independent test on a laptop with a 12 GB RTX 5070 Ti, 64 GB of DDR5 and a gen4 SSD, using ISTA-DASLab's IQ3_XXS quantization: at 43K tokens of context depth the model still generated at 51 tokens per second, versus 23 tokens per second for stock llama.cpp with the same file. Prompt processing showed an even wider gap — about 1,500 tokens per second for a 32K-token document against roughly 100 on llama.cpp. The tester's verdict: "I tried multiple llama.cpp forks and none of them comes close."

Strata's trick is to treat the whole PC as one memory hierarchy rather than fighting the VRAM wall. Qwen3.8-Flash-Next is a mixture-of-experts model with 24,576 specialist networks, of which only about 10 are needed for any given token — around 6 billion of the 125 billion parameters are active at any moment. The GPU keeps the layers used for every token (attention and DeltaNet mixers, routers, shared experts, the output head and the KV cache) plus an adaptive cache of the few thousand most frequently used experts, which it learns as you chat. All 24,576 experts stay pinned in system RAM; when a token needs one the GPU does not have, an x86 CPU kernel computes it with AVX-512 or AVX2 instructions — including a dedicated AVX-512 VNNI kernel written for the Q2_0 format — at the same time the GPU processes its own share, so neither side waits. A 29 GB n-gram lookup table stays on the SSD, with the operating system's page cache prefetching the few small rows each token requires.

On top of the memory-tier design, Strata uses multi-token prediction for speculative decoding: a small built-in helper drafts up to three tokens, one verification pass through the model's 48 layers confirms them, and an average of 2.4 to 3.2 tokens are accepted per pass. Because the large model still rules on every token, output stays bit-identical to standard greedy decoding while arriving 1.6-1.8 times sooner. Long documents are read in chunks of up to 8,192 tokens, which is what pushes prompt ingestion past 1,000 tokens per second.

The installer also supports splitting the model across two or three NVIDIA cards (RTX 20 series or newer, 8 GB or more each), piping prompts through them layer by layer — on an RTX 5080 plus RTX 3090 combination, prompt reading was 18-20 percent faster than on the 5080 alone. An experimental AMD Radeon RX 7900 path exists on Linux via ROCm/HIP. Beyond the base model, the project packages two variants: a Coder build from ISTA-DASLab that strips half the experts while keeping 91 percent of the full model's SWE-bench Verified score and 99 percent of LiveCodeBench, small enough to fit a PC with only 32 GB of RAM; and Swift 1.5, a fine-tune by UkisAI that thinks shorter before answering. Quantized sizes range from 37.6 GB (Q2_0) to 54.8 GB (IQ3_S) of combined RAM and VRAM, with the mid-range IQ2_XS recommended as the default.

The limits are real, and the project is unusually candid about them. Strata serves one request at a time, decodes greedily only, does not reuse the KV cache across conversation turns, and slows down on prefills beyond 128K context. The first start loads 35-55 GB into RAM and can leave the PC unresponsive for one to three minutes while memory is locked and fitted; the model download itself is around 70 GB and needs about 80 GB of disk. Long-context prompt reading also slows as the window fills — the 2,170 tokens-per-second figure applies to 32K-token prompts, dropping at larger sizes.

Strata lands in the middle of a quiet argument about where local inference is going: whether general-purpose engines like llama.cpp and vLLM should keep absorbing every model, or whether the future is a wave of model-specific "one-off" engines that trade generality for large, targeted speedups. Strata — which credits llama.cpp/ggml components and borrows ideas from earlier projects like Splash and ninfer — is a strong data point for the second camp. If a single developer can make a 125-billion-parameter model feel native on a $500 graphics card, the practical floor for private, offline, agent-capable AI keeps dropping — and the audience for that is no longer just enthusiasts, if the repost counters on security researchers' Weibo accounts are anything to go by.

Comments (0)

Log in to join the discussion

Log In

No comments yet