• Home
  • News
  • Blog
  • Releases
  • LLM history
  • Compare LLMs
  • Library
  • About
⌘K
Sign in

A blog and notes on development. The easiest way to reach me is via the social links below.

Contacts
talalaev.misha@gmail.com
Documents
Personal data processing policyPersonal data processing consent

Volundr 32B saves KV cache but is still weaker at agent tasks.

Sh0ny
Sh0ny
5 October 2026
  1. Home
  2. Blog
  3. Volundr 32B saves KV cache but is still weaker at agent tasks.
1 min read

In short

In Agens Volundr 32B Preview, the KV cache is stored in only 18 of the 72 layers, which helps with handling long context. But the memory savings have to be weighed against the lag in agentic benchmarks and the need to run the model on a special build.

Agens Volundr 32B Preview has a strange compromise: the KV cache exists in only 18 of the 72 layers, but the claimed context window reaches 262,000 tokens. If you run models on your own hardware, this could help fit a long context into the available memory.

The remaining layers use linear attention with a fixed-size state or compressed sparse attention. According to the team's measurements, generation speed in BF16 on two 48 GB GPUs stays around 24 tokens per second, consistently from 8K to 128K context. These figures come from the developers themselves, however; there are no independent tests yet.

In some benchmarks, the model outperforms Qwen3.8-27B; in others, it performs at roughly the same level. But Volundr falls behind on agentic tasks: 74.2 versus 79–80 on tau2-bench, and 44 versus 58–64 on a subset of SWE-bench Verified. The team itself calls handling long agent sessions the Preview's main weakness and says the full version will focus on this.

The limitations are significant: for now, the model can only be loaded through Blockway's sglang build; standard sglang and vLLM cannot run it. GGUF and llama.cpp are only promised. In addition, the project is still in training, so the Preview results should not be taken as representative of the final version's characteristics. The code and weights are published under Apache-2.0: project page, BF16 model.

If you run models locally, which matters more to you: a long context with lower memory requirements or more reliable agent performance? Source: LocalLlama.

NewsaillmDevelopment
More AI-tool write-ups on the Telegram channel — short and to the point
Subscribe

Comments

(0)
​