Agents

Cognition Cuts Devin AI Agent Costs by Up to 70%

Cognition has updated its Devin coding agent with intelligent model routing, command batching, and prompt caching to slash operating costs by up to 70% while maintaining benchmark performance.

AlphaSignal3 days agoAgents
Image: AlphaSignal

Cognition has rolled out a major efficiency update for its Devin software engineering agent, targeting the high inference costs associated with long coding sessions. The company reports that operating expenses have dropped by 30% to 40% in Fusion and Normal modes, 15% to 20% in Ultra mode, and up to 70% in Devin Review. To achieve this, Cognition redesigned its agent harness to route tasks to specialized models, batch tool calls, and optimize prompt caching.

Under its "model independence" framework, Devin assigns sub-tasks to specialized LLMs. It routes complex reasoning to Opus 5.5 and GPT-6 Sol, computer interaction to GPT-6 Astra, and supporting tasks to GPT-6 Luna. The system also utilizes Cognition's SWE-2 model, post-trained via reinforcement learning using Moonshot AI's 2.8-trillion-parameter Kimi K3. Available only within Devin, SWE-2 scored 50.0% on the FrontierCode 1.1 Main benchmark, trailing Fable 5.1 by one percentage point at a 64% lower cost. Meanwhile, Devin Fusion scored 68.8 on FrontierCode 1.1 Extended, averaging $0.60 per task.

The updated harness also batches related commands, like formatting code, running a linter, and executing tests. In one example, this reduced a four-turn interaction to two, saving 49% of tokens. Additionally, Cognition restructured Devin's prompts to maximize cache reuse. In a five-turn test, this prompt caching cut tokens processed from scratch by 71%, dropping from 93,000 to 27,000. These savings are amplified by lower provider pricing: Opus 5.5 cache reads cost 60% less than Opus 5, and GPT-6 Sol and Luna cache reads are half the price of their GPT-5.6 predecessors.

For development teams, these updates make deploying autonomous coding agents far more economically viable. However, because FrontierCode is Cognition's proprietary, unaudited benchmark and SWE-2 remains locked behind Devin's closed ecosystem, practitioners should run their own repository trials. Measuring metrics like completion rates, cache-hit rates, and cost per accepted change will help teams verify if these savings translate to real-world codebases.

This is our own summary of reporting by AlphaSignal

More in Agents