Agents

Anthropic Adds Automated Agent Optimization to Claude Code

Anthropic has introduced new evaluation and optimization workflows for Claude Code, allowing developers to systematically improve AI agents without falling victim to benchmark overfitting.

The Neuron3 days agoAgents
Image: The Neuron

Anthropic has launched two new command-line workflows, /claude-api build-eval and /claude-api hillclimb, designed to help developers build rigorous testing suites and automatically optimize their AI agents. The system addresses a common industry pitfall where developers inadvertently optimize prompts for specific test cases without improving real-world performance. By automating the creation of production-representative evaluations and running controlled optimization loops, the tool aims to turn agent development into a disciplined engineering science.

The build-eval tool interviews developers to construct a localized test suite using production transcripts or synthetic data, warning users to optimize for cost or latency if accuracy already exceeds 95 percent. Once a baseline is established, the hillclimb command initiates an optimization loop. It splits the evaluation into a training set and a hidden, held-out test set. Claude then iteratively modifies the agent's prompts, tools, or model configurations, keeping changes only if they improve performance on both datasets. If a modification improves the training score but flatlines on the held-out set, the system reverts the change to prevent overfitting.

Anthropic demonstrated the workflow's efficacy across two internal benchmarks. In a customer-support test of 44 tickets, with 30 used for training and 14 held out, the system started with Opus 4.8 at high effort. Through automated iteration, Claude simplified prompts to reach 87.8 percent accuracy at 1.9 cents per ticket, then swapped the model to Sonnet 5 at low effort to hit 88.9 percent accuracy at roughly 1 cent per ticket. The final optimized configuration achieved 90.5 percent accuracy on the unseen held-out tickets at approximately one-fifth of the original cost. In another test optimizing Anthropic's own Claude API skill, the workflow raised performance from 66 percent to 87.9 percent by round 24, identifying eight missing API features and correcting C# and Java type tables.

For AI practitioners, this development shifts agent design away from trial-and-error prompt engineering toward continuous integration. By measuring the statistical noise floor before making changes and utilizing programmatic or rubric-based LLM judges, developers can reliably distinguish genuine progress from random variance. However, Anthropic notes that human oversight remains critical, as automated hillclimbing will still aggressively optimize toward flawed rubrics or unrepresentative benchmarks if left unchecked.

This is our own summary of reporting by The Neuron

More in Agents