Poolside, a United States startup, has released Laguna S 2.1, an open weight model built for agentic coding, and the story is its size. It is a mixture of experts model with 118 billion total parameters but only about 8 billion active for any given token, which makes it cheap to run, and it is beating systems many times larger on the benchmarks that matter for code.
The numbers are the argument. Laguna S 2.1 scores 78.5 percent on SWE-Bench Multilingual, topping the published leaderboard outright, and 70.2 percent on Terminal-Bench 2.1 with its reasoning mode on. On a long horizon coding benchmark called DeepSWE, it scores 40.4 percent against a top DeepSeek model's 9.0 percent while carrying roughly one sixth the active parameters. It holds its own against models several times its size, including systems from DeepSeek, NVIDIA and Thinking Machines.
The economics follow from the architecture. Because only 8 billion parameters fire per token, the model is cheap to serve, and Poolside has made it free on OpenRouter at shorter context lengths and up to 50 times cheaper than frontier models at the full million token length. For teams whose agents grind through long coding tasks, price per token and context length are the whole decision, and a small open model that scores like a big one changes that math.
The framing is deliberately geopolitical. For most of the past year, the small, cheap, strong open weight model has been a Chinese specialty, with DeepSeek, Qwen and Kimi setting the pace and Western labs mostly shipping large closed systems. Laguna S 2.1 is being pitched as the West's answer, a bet that behavior and training matter more than raw scale, and that an American lab can win on the same efficient open ground the Chinese labs have owned.
The caution is the usual one for a young lab's benchmark table. The leaderboard is Poolside's own compilation, and independent evaluation and real world adoption are what will confirm whether Laguna S 2.1 performs as well in practice as it does on the chart. But the direction is clear and it matters, because the most interesting front in the model race right now is not the largest system, it is the smallest one that can still do the job.
