A 27-billion-parameter model that usually occupies about 54 GB of memory has been squeezed into a 7.89 GB file, and in the one benchmark its publisher cares most about, it came out ahead. Underdog Saluki 27B 1.0, released on October 9 by ConwayResearch on Hugging Face under the Apache 2.0 license, is a two-bit mixed-precision GGUF build of Alibaba's Qwen3.8-27B tuned specifically to preserve tool calling, the capability that turns a chat model into an agent.
The construction is layered. The base is Qwen3.8-27B, a dense 27-billion-parameter model with 64 layers that mixes Gated DeltaNet linear attention with gated attention and supports a 262,144-token context natively. On top of that sits ISTA-DASLab's GSQ-RCO quantization work, which learns accurate low-bit scalar grids per tensor and assigns a quantization type to each tensor under a fixed size budget. ConwayResearch's own pass then shrinks the file further and targets tool calling; the result is tagged IQ2-mix with an importance-matrix flag, and the full recipe for that final pass has not been published.
The headline numbers come from the publisher's own harness and should be read accordingly. On an Underdog Bench of 120 tasks drawn from BFCL v4 and frozen before testing, with thinking mode off and temperature zero, Saluki completed 88 tasks against 84 for the full-size original and 70 for PrismML's Bonsai 2. On 100 parallel tool-call tasks checked with the official BFCL v4 checker, it completed 42 against 35 for the uncompressed model. These are company-published results, not independently verified.
The trade-offs are where you would expect. Competition math is the biggest casualty: AIME 2025 falls from 96.7 to 79.2, roughly 82 percent retention, and multi-step reasoning drops 12 to 18 points. Across nine benchmarks the publisher reports about 96 percent average retention. On SWE-bench Verified, limited to 50 issues, Saluki fixed 30 against 33 for the full model. Each vendor uses its own harness, so the cross-vendor comparisons are not directly comparable.
Deployment is the practical difference. The file runs in stock llama.cpp with a single command, which means it also runs in the applications built on it, and two optional vision projector files, 629 MB and 928 MB, add image input. Two-bit quantization is the most lossy tier available, and structured output is precisely the capability that degrades first under compression, so independent evaluation will matter before anyone ships this in production.
Still, the release is a data point about where local agents are heading. The bottleneck for running useful agents on machines you own has never been conversational quality; it is whether the model can reliably emit valid structured function calls within a memory budget. A 27B-class model that does that from under 8 GB, with no proprietary runtime required, is exactly the configuration that legal, healthcare and air-gapped deployments have been asking for.
Adoption so far is modest but real: roughly 15,000 downloads in the past month and 184 likes at the time of writing, enough to put an anonymous organization on Hugging Face's trending board next to models from far larger labs. Whether the tool-calling edge survives independent testing will decide whether Saluki stays a curiosity or becomes a default local-agent download.
Comments (0)
Log in to join the discussion
Log InNo comments yet