Skip to content
AI.info

The Pulse

Argo-Bench Finds Data Agents Miss the Decision, Not Just the SQL

Argo-Bench evaluates whether AI agents can investigate an enterprise-scale simulated warehouse and make decisions, not just generate SQL. Its results show that models can analyze data yet still miss the objective, while the authors caution

Argo-Bench Finds Data Agents Miss the Decision, Not Just the SQL

AI.info Team ·

TextQL’s October 1 paper tests agents beyond SQL

TextQL researchers submitted Argo-Bench to arXiv on October 1, describing a benchmark for testing whether AI agents can make business decisions after investigating a large company database. Its 210 tasks ask agents to do more than return a query: they analyze records and file decisions such as banning accounts, allocating courier incentives or issuing back pay.

The benchmark simulates a New York food-delivery platform with 81 million orders in 2024. The paper describes the warehouse used in its tests as having 235 tables and 7.5 billion rows, modeled on Oracle E-Business Suite. The public warehouse released separately has 7.54 billion rows; the paper says the private test world has 7.49 billion. Tasks cover fraud, marketplace operations, financial planning, accounting and growth.

Agents act on facts hidden from the warehouse

The simulator keeps its ground-truth state separate from the warehouse agents inspect. That requires a model to reconstruct what happened by tracing information across tables before filing an action. Graders score those actions against the simulator’s hidden state, rather than judging only the agent’s written answer. Every task also has an executable reference solution that demonstrates how it can be solved using the warehouse.

Tasks include actions, forecasts, budgets, reported figures and dashboard data sources. In 99 tasks, agents receive only a partial view of the year. The paper says those partial views are used mostly for forecasting, not exclusively; its reported forecast results cover 72 tasks. Each agent ran in a sandbox with Python tools and no outbound internet access.

Claude Opus 5.5 leads the model test

Across 14 frontier and open-weight models, Claude Opus 5.5 led overall, averaging 59.5 out of 100. It scored at least 95—the paper’s threshold for a solved task—on 34.8% of tasks. Table 9’s 95% confidence intervals extend as much as 9.3 points above a score estimate and 8.5 percentage points above a solved-task estimate; the intervals reflect the choice of tasks, not run-to-run variation.

A courier-bonus task illustrates how agents can analyze data carefully yet miss the business objective. The paper’s authors—Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand and Joseph J Ma—write:

“The platform pays for quests to avoid surge pay; however, under the simulator’s response model, the plan loses $86,281 where a uniform cut would save $0.40 million, and it scores 0.”

GPT-6 Astra identified where bonuses had been withheld at random and estimated how courier supply responded, but its proposed cuts lost $86,281 under the simulator’s response model. Claude Opus 5.5 instead accounted for bonuses’ role in replacing surge pay and saved $3.09 million of a possible $3.12 million. In forecasting, nominal 80% intervals contained the eventual value in just 44.8% of 4,553 forecast series across 72 tasks.

The benchmark measures a simulation, not a real company

Argo-Bench does not directly predict performance on a company’s warehouse. The paper says the simulation covers one city and one year and supports one ERP format. It also identifies weakly grounded assumptions in its simulated membership program and cautions that calibration to aggregate targets does not guarantee realistic edge cases. The authors say the benchmark compares data agents rather than estimating how they would perform at a real company.

The public warehouse uses a separate simulator seed from the private world used for official scoring, so its entities and answer keys differ. Researchers can download the warehouse from Hugging Face under CC BY 4.0. The GitHub repository provides the tasks, reference agent and evaluation code under Apache 2.0.

Sources

How it is written

Drafted with AI models and checked under AI.info’s standard: the source is opened, every fact is cross-checked against at least two independent sources, and an adversarial reading looks for what is wrong. The editor answers for what is published. The standard

Explore

More articles