Skip to content
AI.info

The Pulse

Surge AI’s DAYJOB Tests Whether Agents Can Make the Right Call

Surge AI released DAYJOB, a benchmark suite for professional agents in healthcare and finance, on September 24, 2026. The company says the top-scoring models fully passed fewer than one in four assignments in either field.

AI.info Team ·

Surge AI says its new benchmark asks agents to make the calls that conventional work tests may already spell out. The company’s critique of OpenAI’s GDPval—that its tasks often prescribe the steps and measure routine production—sets up a direct disagreement over what counts as professional capability. DAYJOB, announced September 24, is Surge’s attempt to test the judgment that comes before a finished document.

Surge’s challenge to GDPval

The first DAYJOB benchmarks cover healthcare and finance. Surge says the suite contains 130 expert assignments, built around brief workplace requests and large collections of records, reports, spreadsheets and emails. Instead of telling an agent exactly what to do, the task may ask whether a report is safe to send or whether a patient can be discharged; the agent must identify relevant evidence and decide what it means.

Surge contrasts DAYJOB’s average prompt lengths—81 words in finance and 49 in healthcare—with 337 words for GDPval tasks. It also says DAYJOB assignments provide far more supporting files: an average of 25.7 in finance and 19.8 in healthcare, compared with 1.2 for GDPval. Those are Surge’s comparisons and its interpretation of GDPval’s limits; they do not establish that the older benchmark fails to measure every kind of professional skill.

A unit error swells to $25.7 million

One finance assignment asks an agent to inspect a fund’s reports before they go to a new client. The supporting materials include policy documents, vendor price files, system reports, broker quotes and emails. Surge says a position was priced from a Bloomberg screenshot showing South African cents, but its value was entered as rand—making a roughly $260,000 holding appear to be worth $26 million.

GPT-5.6 caught two other pricing-source errors, worth a net $132,000, and recommended delaying the reports. It missed the much larger unit error. Across 66 runs in 22 model configurations, Surge says, only two caught the cents-versus-rand mistake; both were Claude Opus 5 runs. The example tests whether an agent can trace a number to its source and challenge a plausible-looking result, rather than simply complete calculations.

The healthcare case turns on a diagnosis

In a healthcare task, a clinician asks an agent to review lab results before discharging a patient with facial weakness described in the handoff as “textbook Bell’s Palsy.” Other records show neck pain after heavy lifting, difficulty speaking and an incomplete facial examination—details Surge says should prompt further assessment before discharge.

Surge reports that GPT-5.6 recognized the possibility of a vascular cause in its reasoning but decided not to question the diagnosis unless necessary. A Grok run reviewing the same records challenged the plan and recommended holding discharge for emergency evaluation and vascular imaging. The contrasting responses illustrate the benchmark’s stated focus: whether an agent notices when evidence should change the requested course of action.

A full pass requires every rubric criterion

Surge reports that the strongest model fully passed 24.7% of DAYJOB: Healthcare assignments and 23.9% of Finance assignments. Those figures measure complete task passes, not whether a model earned partial credit: the benchmark’s public dataset cards define success as satisfying 100% of a task’s rubric criteria. The distinction matters when reading a low pass rate as evidence about overall performance.

Surge says experts with direct experience in the fields created the tasks and that each assignment went through three review layers for correctness, completeness and solvability. The company has published task materials on Hugging Face for Healthcare and Hugging Face for Finance, along with a GitHub evaluation harness. The results come from a benchmark made by Surge, so they offer a test of its chosen definition of workplace judgment—not an independent measurement of how often agents fail on the job.

Source

Explore

More articles