Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling






Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers, that means a 27B-class agent model that runs in stock llama.cpp.

TL;DR

  • Size: 27B dense parameters. 7.89 GB GGUF versus 54 GB for BF16.
  • Runs on: stock llama.cpp and apps built on it, with full GPU offload. Optional 629 MB or 928 MB vision add-on.
  • Performance: 96% average retention across 9 benchmarks versus full Qwen3.8-27B.
  • Best: Parallel tool calls, 42 versus 35 for the full model (120% retention).
  • Worst: AIME 2025, 79.2 versus 96.7 (about 82% retention).
  • Bottom line:
    • Best: beats the 54 GB original at tool calling in a sub-8 GB file.
    • Worst: competition math and multi-step reasoning drop 12 to 18 points.

What is Underdog Saluki 27B?

Saluki 27B is a 2-bit, mixed-precision GGUF of Qwen3.8-27B built for local agents. It stacks 3 layers of work:

  • The base is Qwen3.8-27B, a dense 27B model from the Qwen team. It has 64 layers, mixes Gated DeltaNet linear attention with gated attention, and supports 262,144 tokens natively.
  • The second layer is ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF. GSQ learns accurate low-bit scalar grids per tensor. RCO assigns a quantization type to each tensor under a fixed size budget. ISTA’s smallest file, IQ2_XS, is 8.4 GB at 2.50 bits per weight.
  • The third layer is Underdog’s own pass. It shrank the file to 7.89 GB and targeted tool calling. The file is named IQ2-mix and carries an imatrix tag. Underdog has not published the full recipe for this pass.

How does Saluki perform on benchmarks?

Underdog splits its results into 2 groups.

The first group ran both models in the same harness:

  • Underdog Bench: 120 tasks from BFCL v4, frozen before testing. Thinking off, temperature 0. Saluki scores 88, the full model 84, and PrismML’s Bonsai 2 scores 70.
  • Parallel tool calls: 100 BFCL v4 parallel tasks with the official checker. Saluki 42, full model 35.
  • SWE-bench Verified: 50 issues. Saluki fixes 30, the full model 33.

The second group compares Saluki with public full-size scores:

BenchmarkSaluki 27BQwen3.8-27B (public)
IFEval (prompt-loose)93.591.5
IFBench (prompt-loose)72.771.0
MBPP+78.083.9
MuSR67.579.6
AIME 2025 (avg@4)79.296.7
AIME 2026 (avg@4)80.094.6



Source link

  • Related Posts

    Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work

    Google Cloud has introduced the Google Cloud Gemini agent, a single agent for enterprise work. The Gemini agent is a cloud-hosted agent from Google Cloud that answers questions, does knowledge…

    Google Research RRSI Guide: Mastering Self-Improving AI Agents

    In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen…

    Leave a Reply

    Your email address will not be published. Required fields are marked *