Agents

Atria Dawn project shows humans still guide AI development

A study of the development of the Atria Dawn Preview model reveals that while AI agents perform the bulk of execution tasks, human engineers still make the vast majority of critical decisions.

The Decoder4 days agoAgents
Image: The Decoder

A research team including members from China's Fudan University analyzed their own collaborative workflow while building Atria Dawn Preview, a 744-billion-parameter mixture-of-experts language model designed for research and engineering. The model leads on five of 16 benchmarks, including AutomationBench, CyberGym, and MLE-Bench Lite, though it trails in GDPval and SWE-Bench Pro. By reviewing over 700 task logs from 56 participants, the researchers mapped out exactly how much work AI agents performed compared to their human supervisors.

The data shows that AI was utilized in 96.5 percent of the analyzed tasks. Over a four-week period, the median ratio of agent actions to human inputs grew from 11 to 28.5. This increase did not represent true autonomy, however, as humans still made 85.5 percent of decisions regarding methods and parameters, while AI agents made only 9.2 percent. The most frequent workflow pattern, occurring in 55.4 percent of cases, involved the AI proposing options for the human to select. Furthermore, humans determined the final goals and scope in 93.4 percent of all tasks.

Interestingly, participants reported that 151 of the 455 completed AI-assisted tasks—roughly one-third—would have been completely impossible to attempt without AI assistance. When technical difficulties arose across 588 recorded instances, human intervention was required to move forward 76 percent of the time. This intervention rarely involved manual coding; instead, humans provided context or clarified requirements in 35.2 percent of cases, or diagnosed issues in 34.7 percent of cases. Once feedback was given, the AI successfully executed the revisions on its own 75.4 percent of the time.

For AI practitioners, these findings suggest that the bottleneck in model development has shifted from execution to high-level judgment. While companies like Anthropic claim to automate research directions and OpenAI deploys GPT-5.6 Sol, this study aligns with findings from Princeton and the UK AI Security Institute suggesting that human oversight remains indispensable. Rather than replacing engineers, agents act as force multipliers, allowing developers to focus on system design and troubleshooting while delegating the tedious labor of code generation and execution.

This is our own summary of reporting by The Decoder

More in Agents