Research
World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems
Overview Research area: Enterprise AI agents, LLM evaluation benchmarks, and world modeling for partially observable symbolic systems. Technical level: Advanced. The paper is written for readers comfo

- arXiv
- 2601.22130
- Published
- 2026-01-29
- Authors
- Lakshya Gupta, Litao Li, Yizhe Liu, Sriram Ganapathi Subramanian, Kaheer Suleman, Zichen Zhang, Haoye Lu, Sumit Pasupalak
AI summary
Overview
Research area: Enterprise AI agents, LLM evaluation benchmarks, and world modeling for partially observable symbolic systems.
Technical level: Advanced. The paper is written for readers comfortable with partially observable Markov decision processes (POMDPs), tool-calling / MCP agents, and reinforcement-learning concepts such as forward and inverse dynamics. Its high-level conclusions are accessible, but the environment specification and metrics are formal.
Scope: The paper introduces WoW, a ServiceNow-based enterprise environment containing hidden workflows and business rules, together with WoW-bench, a 234-task benchmark that measures whether frontier LLMs can act reliably and predict the cascading side effects of their actions.
What This Paper Is About
LLM agents are increasingly deployed against enterprise systems, but those systems hide workflows and business rules that silently change database state when an action is taken. A single API call can trigger cascading updates across multiple tables, so an action that looks locally valid can violate a constraint the agent never sees.
Existing enterprise benchmarks mostly test surface-level task completion — UI navigation, retrieval, instruction following — and their tasks can usually be solved from observed information alone, which effectively makes them fully observable. This paper builds a workflow-heavy environment and a benchmark that forces agents to confront the observability gap, then measures how badly they fail and how much of the failure is a lack of information versus a lack of reasoning.
Key Contributions
-
WoW, a high-fidelity enterprise environment. A ServiceNow-based system with workflows and business rules embedded in it, covering management sub-domains for user, incident, asset, knowledge base, catalog, and expense. Agents act through MCP tools and observe the world through either tool responses or oracle table audit logs.
-
WoW-bench, a 234-task benchmark. The benchmark is split into four categories: 67 action prediction tasks, 67 audit prediction tasks, 50 constraint understanding tasks, and 50 agentic tasks. It is the first enterprise benchmark designed to evaluate LLMs both as enterprise world models and as agents, and it explicitly incorporates workflow effects.
-
An observation-space ablation. The same tasks are run under two observation functions — tool response only versus tool response augmented with table audits (differential state changes) — to isolate how much reliability depends on state visibility.
-
A diagnostic error taxonomy. The evaluation is followed by an analysis organizing failures into three gaps: the Representation Gap, the Dynamics Gap, and the Causal Gap, with implications for dynamics-aware agent architectures.
Main Findings
-
Constraint reliability collapses without state visibility. GPT-5.1 reaches only 2% Task Success Rate Under Constraint (TSRUC) with tool-response observations, improving to 14% with table audits. Under tool observations, TSRUC is 6% for Gemini-3-Pro, 4% for Sonnet-4.5, and 8% for Opus-4.5; under audit observations it is 16%, 30%, and 14% respectively.
-
Task success rate is much higher than success under constraint. Under audit observations, TSR is 32% for GPT-5.1, 42% for Gemini-3-Pro, 58% for Sonnet-4.5, and 36% for Opus-4.5. Under tool observations, TSR is 22%, 38%, 32%, and 26%. Agents complete goals while silently violating constraints.
-
Audit logs provide a large but bounded uplift. Using audit logs as the observation increases task success rate by at most 7x, and accuracy can increase by at most 10x for constraint understanding tasks. Even with this oracle-level visibility, absolute success under constraint remains low.
-
Cost does not track reliability. Average cost per task under audit observations ranges from $0.41 (GPT-5.1) to $5.00 (Opus-4.5), with Sonnet-4.5 at $3.60 and Gemini-3-Pro at $0.64. Under tool observations the range is $0.14 to $3.08. The paper states Opus-4.5 is the most expensive while being on par with Gemini-3 for TSRUC.
-
World modeling accuracy is low. For both audit prediction (forward dynamics) and action prediction (inverse dynamics), LLMs score below 30%. The strongest models achieve near-zero accuracy in full state prediction in audit prediction tasks.
-
Models under-predict state changes. Models consistently miss side effects, such as a Create Incident action silently updating the
metric_instancetable, suggesting they rely on semantic similarity rather than a learned transition model. -
Symbolic grounding is the dominant error source. Conflating human-readable names with unique identifiers (for example, predicting username versus
sys_id) accounts for 73.5% of errors. -
Failures are structural, not stochastic. The paper attributes them to three gaps: models treat entities as text tokens rather than nodes in a relational graph (Representation Gap); they lack an explicit forward transition model P(s_{t+1} | s_t, a_t) (Dynamics Gap); and they behave as greedy planners that optimize the immediate action while ignoring multi-hop cascades such as Action → Workflow A → Workflow B → Violation (Causal Gap).
Methodology in Plain English
The researchers took a free ServiceNow developer instance — a real enterprise platform — and configured it so that its ordinary automation machinery would be doing the work during evaluation. ServiceNow exerts its effects through two mechanisms: business rules, which are immediate, atomic, database-driven logic that can set one column based on other values in the same table; and workflows, which orchestrate multi-step, multi-system processes. Agents never see these definitions; they only see what the platform tells them after each call.
Tasks are modeled as a POMDP. The user query supplies the task and constraint description. The state is the entire underlying relational database, which is treated as intractably large and never fully observable. Actions are MCP tool calls made of a tool name plus free-form parameter key-value pairs — the paper notes this action space is combinatorially large and demands high precision. The paper extended an existing ServiceNow MCP server (Echelon Lab MCP) to 108 tools supporting create, read, update, and delete operations.
The key experimental lever is the observation function. Under the standard setting, the agent sees only the tool's immediate output — success messages, error codes, fetched records — which hides all downstream side effects. Under the oracle setting, the tool response is augmented with table audits: a structured list of differential database changes, each entry a tuple of (TableName, ColumnName, OldValue, NewValue) covering every update during the transition, including those caused by hidden asynchronous workflows. The audit setting is still partially observable, because the workflow specifications themselves are never exposed.
Task construction used domain experts. For constraint understanding, experts hand-wrote 10 realistic constraints (for example, "A user cannot be assigned more than 3 active incidents at a time" and "Flagged articles should not be published"), then designed actions with deliberate cascading effects, producing 10 templates later perturbed into 50 trajectories. For agentic task completion, 10 task templates were derived from those constraints, converted to task descriptions by an LLM and verified by humans, with 5 permutations each for 50 tasks and an average of 13 actions per task. For action and audit prediction, the paper uses a Tool-Dependency Graph Sampling technique to build connected multi-hop trajectories where tool outputs feed into subsequent inputs, rather than random disjoint samples. Cleanup functions delete task-specific data so runs can be reproduced repeatedly.
Evaluation uses Task Success Rate and Task Success Rate Under Constraint (which multiplies goal satisfaction by one minus the violation indicator), plus average cost per task in US dollars. Constraint understanding requires an exact match identifying both the violated constraint and the responsible action. Dynamics modeling is scored with Intersection over Union for audit prediction, requiring exact equality of each (Table, Column, OldVal, NewVal) tuple, and with tool name accuracy and full action accuracy for action prediction.
Why This Matters
The paper argues that the bottleneck for enterprise autonomy is not instruction following, context length, or tool selection, but dynamics understanding. Benchmarks like WorkArena++, CRMArena-Pro, SCUBA, ST-WebAgentBench, WorkBench, τ²-bench, MCPToolBench++, and MCP-Universe each cover some of realistic environments, enterprise tasks, constraint following, world model evaluation, or complex workflows — WoW-bench's comparison table claims all five, where no prior benchmark claims more than three. If that framing holds, evaluation should shift from surface-level completion toward deep constraint reliability and hidden-transition prediction.
Real-world applications:
- IT service management and helpdesk automation. Agents assigning assets, roles, and incidents in ITSM platforms can trigger clearance, SLA, and escalation workflows that silently undo or conflict with their work.
- CRM and revenue operations. Read-write operations across accounts, opportunities, and approval chains involve multi-hop dependencies that read-only benchmarks like CRMArena-Pro do not exercise.
- Compliance and audit tooling. The finding that TSRUC is near zero while TSR stays moderate means systems could pass functional tests while quietly generating violations — relevant to regulated deployments.
- Agent architecture and infrastructure. The results motivate building persistent structured state representations and learned transition models instead of relying on prompt engineering or larger context windows.
Industry relevance is direct: enterprises run domain-engineered systems with system-specific requirements, and the paper notes that because such environments are unlikely to appear in large-scale pre-training, general-domain capability does not transfer. The economic argument is also concrete — cost per task varied by more than an order of magnitude across models without a corresponding reliability gain.
Future Directions
- Learning system dynamics explicitly. The paper calls for Model-Based Reinforcement Learning agents that learn a predictive model of hidden workflows and simulate the audit log before execution, using prediction-reality discrepancies to update the internal model.
- Structured, persistent state representations. Agents need entity abstractions that track values, statuses, and relationships independently of how entities happen to be mentioned in the conversation history, to address the Representation Gap that produced 73.5% of errors.
- Active epistemic strategies. Instead of assuming a static world, agents could issue probe actions to test hypotheses about workflow triggers and boundary conditions, learning the environment's physics rather than waiting to observe it.
- Benchmark expansion and training use. The paper acknowledges that expert curation makes expansion relatively expensive, and that the current dynamics cover only a subset of workflows in a single enterprise system. It also notes WoW is intentionally an evaluation benchmark rather than a training environment, and that demonstrating gains from training-time interaction is left to future work despite the fully interactive design.
Target Audience
Researchers and engineers working on LLM agents, tool use, and agent evaluation; enterprise AI practitioners deploying agents into platforms such as ServiceNow or CRM systems; and reinforcement learning researchers interested in world models and partially observable symbolic environments. The paper also suits benchmark designers who want a concrete example of how to construct tasks whose difficulty comes from hidden system dynamics rather than from task length or UI complexity.
Authors’ abstract
Frontier large language models (LLMs) excel as autonomous agents in many domains, yet they remain untested in complex enterprise systems where hidden workflows create cascading effects across interconnected databases. Existing enterprise benchmarks evaluate surface-level agentic task completion similar to general consumer benchmarks, ignoring true challenges in enterprises, such as limited observability, large database state, and hidden workflows with cascading side effects. We introduce World of Workflows (WoW), a realistic ServiceNow-based environment incorporating 4,000+ business rules and 55 active workflows embedded in the system, alongside WoW-bench, a benchmark of 234 tasks evaluating constrained agentic task completion and enterprise dynamics modeling capabilities. We reveal two major takeaways: (1) Frontier LLMs suffer from dynamics blindness, consistently failing to predict the invisible, cascading side effects of their actions, which leads to silent constraint violations, and (2) reliability in opaque systems requires grounded world modeling, where agents must mentally simulate hidden state transitions to bridge the observability gap when high-fidelity feedback is unavailable. For reliable and useful enterprise agents, WoW motivates a new paradigm to explicitly learn system dynamics. We release our GitHub for setting up and evaluating WoW.