Research
Atria Dawn: The Dawn of Agentic Superintelligence
Overview Research area: Agentic AI — foundation language models that use tools over sustained tasks, plus an empirical study of human–AI collaboration inside an AI research-and-development project. Te

- arXiv
- 2609.15818
- Published
- 2026-09-14
- Authors
- Honglin Guo, Tao Gui, Kun Cai, Haodong Chen, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Zerun Ma, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Penghao Zhao, Ruojun Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou, Qiye Cai, Xinmeng Che, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan, Zixin Shang, Bing Shao, Zhuohui Sheng, Jiayang Shi, Yang Shu, Aierpanjiang Simayi, Sirui Song, Yuxiao Song, Zhe Sun, Zhichao Sun, Wenzhe Tan, Wenhui Tian, Zhongbo Tian, Hanchen Wang, Pengbo Wang, Rui Wang, Yiding Wang, Yuhui Wang, Zhiheng Xi, Caijun Xu, Chao Xu, Yongfeng Xu, Xiaolei Yang, Zhixiong Yang, Qian Yao, Shihong Yi, Yuankai Ying, Jia Yu, Dingbo Yuan, Hao Yuan, Junjie Yuan, Bo Zhang, Caixian Zhang, Qiuyinzhe Zhang, Jiyuan Zhao, Ying Zhao, Pujun Zheng, Xiaoxue Zhong, Xiaohao Zhou, Xinyu Zhou, Guanru Zhu, Yulun Zhu, Yaojie Lu, Tao Ji, Hongyu Lin, Yutao Zhu, Pengfei Cao, Guoxiu He, Xianpei Han, Ben He, Zhicheng Dou, Kang Liu, Qi Zhang, Le Sun, Jun Zhao, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Bowen Zhou
AI summary
Overview
Research area: Agentic AI — foundation language models that use tools over sustained tasks, plus an empirical study of human–AI collaboration inside an AI research-and-development project.
Technical level: Advanced. The paper combines a large-scale agentic training pipeline built on a 744-billion-parameter mixture-of-experts foundation model with a 16-benchmark evaluation and a mixed quantitative/qualitative study of R&D workflows.
Scope: The paper introduces Atria Dawn Preview, an agentic language model for research and engineering workflows trained through a Verifiable Experience Pipeline, and uses its own development process as a case study of how work is divided between agents and human researchers.
What This Paper Is About
As language-model agents take on sustained tool use in software engineering, document reasoning, and research, they increasingly participate in building the AI systems that come after them. Task performance alone cannot answer who, in a real model-development project, chooses which problems are worth pursuing, which proposed methods are adopted, and how the next round of effort is directed. The paper addresses this by presenting Atria Dawn Preview as a capable agentic model and by analyzing 769 task records from 56 participants (plus agent logs) to describe how responsibility is actually split between humans and agents.
Key Contributions
-
Atria Dawn Preview, a foundation agentic language model built on a 744-billion-parameter mixture-of-experts foundation model (Z.ai, 2026), designed for scientific research and engineering workflows with the aim of expanding the frontier of agent productivity in the real world.
-
The Verifiable Experience Pipeline, a training recipe that connects tool-mediated interactions to executable environments and externally verified outcomes. Each training task is tied to a real execution environment, and outcomes are checked with external signals such as tests, metrics, file state, geometric structure, source evidence, or human-defined criteria.
-
A 16-benchmark evaluation spanning tool use, search and research, workspace productivity, professional tasks, software engineering, terminal tasks, ML engineering, and cybersecurity, comparing Atria Dawn Preview against six models: DeepSeek V4 Pro 0813, KIMI K3, Qwen 3.8 Max, GLM 5.3, GPT 5.6 sol, and Claude Opus 5. Atria Dawn achieves the highest reported score on five benchmarks and the second-highest on three others.
-
An empirical study of human–AI collaboration in AI R&D, analyzing 769 task records from 56 participants alongside agent logs, covering task feasibility, the attribution of proposals versus final choices, and how work resumes after difficulties.
Main Findings
-
Leading results on five benchmarks. Atria Dawn Preview records the highest reported score on AutomationBench (53.8), BFCL v4 (77.0), DeepSearchQA (96.0), BrowseComp (92.5), and CyberGym (86.5). Its margins are 4.1 points over the runner-up Qwen 3.8 Max on AutomationBench and 2.9 points over GLM 5.3 on BFCL v4; on CyberGym it exceeds GLM 5.3 (84.5) by 2.0 points.
-
Second place, narrowly, on three others. SkillsBench (66.4), Workspace-Bench (65.0), and Workspace-Bench-Lite (68.2) trail the leading scores by only 0.3, 0.8, and 1.9 points respectively.
-
Competitive results across the remaining benchmarks. WideSearch 81.9 (matching Qwen 3.8 Max, ahead of KIMI K3 at 79.6, 1.4 points behind GPT-5.6 sol at 83.3); DeepResearch Bench II 51.1 (KIMI K3 51.3; ahead of DeepSeek V4 Pro 46.6, Qwen 3.8 Max 49.2, GPT-5.6 sol 50.7); τ³-Bench Banking 41.2 (above KIMI K3 37.1 and GLM 5.3 40.2); GDPval 1583 (ahead of DeepSeek V4 Pro at 1517); JobBench 50.3 (ahead of GPT-5.6 sol at 45.4); MLE-bench Lite HumanRank 86.2 (above KIMI K3 85.8, Qwen 3.8 Max 81.3, GLM 5.3 80.8); SWE-bench Pro 59.6 (DeepSeek V4 Pro 58.3, GLM 5.3 60.3); Terminal-Bench 2.1 78.3 (DeepSeek V4 Pro 78.7).
-
AI use is near-universal in this project. Of 739 tasks with a clear response about AI use, 713 involved AI, or 96.5%.
-
Delegation rose sharply over a four-week window. For the same cohort of 22 participants, the daily median ratio of agent actions to human prompts rose from 11.0 to 28.5 between August 7 and September 4, with valid data for 21–22 participants each day. The interquartile range widened over the same window, so the shift is uneven.
-
Roughly a third of AI-assisted tasks were rated infeasible without AI. Among 455 completed AI-assisted tasks with a usable response, 151 (33.2%) were reported as infeasible without AI, holding scope, quality requirements, and other resources fixed. These 151 tasks came from 27 of the 56 participants.
-
Humans propose less often than they select. For methods and parameters, "AI proposes, human selects" was the most common pattern at 55.4% of decisions. Humans made the final choice in 85.5% of method-or-parameter decisions, versus 9.2% for AI. Human final selection was also prevalent for goals or scope (93.4%) and acceptance criteria (81.9%). AI's share of proposals ranged from 16.9% for goals or scope to 55.4% for methods or parameters, while its share of final decisions stayed between 6.1% and 9.2%.
-
Dependence on AI does not mean AI decides. In the subset of 151 tasks rated infeasible without AI, humans selected the final goal in 144 cases (95.4%), a higher rate than in general project decisions.
-
Difficulties were resolved mostly through human help, but that help was informational. Among 588 tasks with a recorded difficulty and response, 76.0% moved forward through human intervention, agents recovered on their own in 23.0%, and only 1.0% were left unresolved. Human help (447 tasks in total) was most often adding context or clarifying requirements (35.2%) or diagnosing the issue or changing the method (34.7%); partial edits were 3.2% and takeover 0.7%. The paper states human help was about eighteen times more likely to change what the agent knew than to change who did the work.
-
Outputs were overwhelmingly adopted and revised by the agent. Of 627 tasks with a known disposition for the main AI output, 354 (56.5%) involved substantive revisions, only 1.0% were not adopted, 3.2% were used for ideas alone, and 95.9% entered the deliverable in some form. Within the 354 substantively revised tasks, the AI made the revisions after human feedback in 75.4% of cases, versus 19.2% where a human edited directly.
-
Demonstrated case capabilities. With web search disabled, Atria Dawn processed more than 100 GB of weather data, implemented a vision-transformer-based network with more than 0.4 billion parameters, and trained it for 45,000 steps to model 69 meteorological variables. In Gated Delta Network decode optimization it switched from Triton to native CUDA, and development measurements yielded a 1.46× ratio between summed baseline and candidate latencies over seven representative batch sizes; 52 of 54 formal workloads had passed, with no complete formal mean available. A MiniOS build from an empty workspace took approximately 20 minutes and included a serial shell, disk access, a persistent filesystem, and an interpreter, with persistence and autorun checked across two QEMU sessions.
Methodology in Plain English
The training approach connects every task to a real environment rather than to synthetic prompts. The agent observes the state, calls tools, produces intermediate artifacts, and adapts to feedback; its final outcome is then checked by an external signal — executable tests, experiment metrics, file and application state, geometric checks, or source support. Trajectory curation removes incomplete, contradictory, duplicate, or behaviorally invalid examples, keeping only the links between tasks, trajectories, artifacts, and verification evidence. Recurring failures such as ineffective tool selection, incomplete verification, missing evidence, and unsuccessful recovery feed back into constructing more tasks and refining environments. Longer workflows additionally record job launches, log and artifact inspection, and revision of execution plans.
Evaluation uses 16 benchmarks with published protocols. For example, BFCL v4 uses the official harness at a specified commit with 5,217 cases, 16 concurrent workers, temperature 0.001, a 120-second timeout per request, at most two attempts, and official weighting of 10% Non-Live, 10% Live, 10% Irrelevance, 30% Multi-Turn, and 40% Agentic. SkillsBench uses OpenHands within the AgentCompass framework on a self-contained subset of 79 tasks, excluding multimodal tasks, averaged over three runs. τ³-Bench Banking uses 97 Banking Knowledge tasks, 33 concurrent workers, a random seed of 300, and an alltools retrieval configuration with a 1,800-second simulation timeout and a 2,400-second outer task timeout.
For the collaboration study, the team combined task records with agent logs. Participants estimated whether their own part of a completed task could have been done without AI under fixed scope, quality requirements, and other resources; they also attributed who proposed and who selected options for goals or scope, methods or parameters, and acceptance criteria, and recalled the most consequential difficulty in each task.
Why This Matters
The paper argues that progress toward more autonomous AI research requires advancing both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development. It also distinguishes time saved from work made possible: an evaluation that measures only hours saved, the authors write, understates what changed.
Real-world applications described in the paper:
- Scientific research and ML engineering, including a weather-forecasting workflow over more than 100 GB of data and a 45,000-step training run, plus Gated Delta Network decode optimization.
- Software engineering and terminal work, illustrated by building MiniOS from an empty workspace in approximately 20 minutes with checks across two QEMU sessions.
- Cybersecurity, where in an authorized, isolated test environment the model inspected a website, validated an injection flaw, repaired the underlying issue, and re-tested the affected paths.
- Professional deliverables and design artifacts, including reports on clean-energy siting, healthcare investment, and semiconductor capital allocation, generated presentations, and CAD assemblies for a four-cylinder engine, a humanoid robot, a robotic joint module, and a rocket.
Industry relevance: For organizations building or deploying agentic models, the benchmark table provides a comparative read on frontier performance across tool use, search, workspace productivity, software and terminal engineering, ML engineering, and security. The R&D study is directly relevant to teams running agents through autonomous modes (the paper notes use of Codex's yolo mode and Claude Code's skip-permissions mode), because it documents where authority actually sat, how often humans had to intervene, and how many agent actions each human judgment decision now propagates through.
Future Directions
-
Generating and evaluating diverse research directions. How can AI keep producing diverse directions and assess their potential value before results are available? The paper recommends establishing communication channels between different working agents, noting that new pipelines and workflows still came from researchers prompting agents toward them.
-
Turning experience into intrinsic capability. AI must recursively transform experience accumulated over long exploration iterations into intrinsic research capability. The paper observes that what an agent gained from a failed run stayed in that session or in its records while the researcher carried the lesson forward, and suggests test-time training may prove valuable and that current architectures separating memory from capability may require fundamental redesign.
-
Sustaining meaningful human oversight. As AI reveals possibilities beyond human imagination, how can humans still formulate meaningful goals and maintain effective oversight? The paper warns that when researchers cannot assess even final outcomes of AI work, oversight becomes impossible, and that human involvement could become a placebo rather than a substantive contribution.
-
Protocols for authority allocation and alignment. Clear guidelines are needed for authority allocation and alignment standards, including whose intentions shape alignment objectives, who accepts the associated risks, and who remains accountable. The paper argues each advance in AI capability should be matched by proportional investments in alignment and oversight.
-
Measurement of research capability itself. The paper states that progress on general benchmarks does not reveal research capabilities, and calls for tasks designed specifically to assess research capability.
Target Audience
This paper is most useful to researchers and engineers working on agentic language models and agent training pipelines, to AI R&D leads and platform teams deciding how to deploy coding and research agents, and to safety and governance researchers interested in authority allocation and human oversight. Readers studying human–AI collaboration, human-computer interaction, or the organization of AI research will find the 769-record study and its breakdown of proposal versus selection roles particularly relevant.
Authors’ abstract
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.