Research
TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
Overview Research area: Natural Language Processing, specifically tool-calling agents and evaluation benchmarks for Large Language Models (LLMs). Technical level: Advanced. Scope: This paper introduce

- arXiv
- 2510.19286
- Published
- 2025-10-22
- Authors
- Reza Esfandiarpoor, Vishwas Suryanarayanan, Stephen H. Bach, Vishal Chowdhary, Anthony Aue
AI summary
Overview
Research area: Natural Language Processing, specifically tool-calling agents and evaluation benchmarks for Large Language Models (LLMs).
Technical level: Advanced.
Scope: This paper introduces TheMCPCompany, a benchmark built from real-world REST APIs (packaged as MCP servers containing over 18,000 tools) for evaluating how well LLM agents can select and combine task-specific tools instead of relying on web browsers.
What This Paper Is About
Since the Model Context Protocol (MCP) was introduced, the number of tools available to LLMs has grown substantially, and these task-specific tool sets can be easier to build and maintain than graphical interfaces. Despite this, most general-purpose agents still interact with the world primarily through web browsers. This paper asks whether agents can instead solve real tasks by calling the right tools from large, real-world tool collections, and it builds a benchmark to measure that ability under both ideal and realistic tool-retrieval conditions.
Key Contributions
- A new benchmark, TheMCPCompany, for evaluating tool-calling agents on tasks that require interacting with a variety of real-world services.
- A large-scale tool environment: MCP servers derived from the REST APIs of those services, exposing over 18,000 tools in total.
- Manually annotated ground-truth tools for each task, used as an upper-bound condition to isolate the value of correct tool selection from the difficulty of finding the right tool.
- An empirical comparison of two regimes — agents given the correct tools versus agents that must retrieve tools themselves — including a comparison against browser-based agents.
Main Findings
- Tools beat browsers when retrieval is imperfect: All models tested with tool retrieval performed similarly to or better than browser-based agents.
- Ground truth shows the ceiling: Using the manually annotated ground-truth tools, the authors demonstrate that tool-calling agents can both improve performance and reduce costs, assuming perfect tool retrieval.
- Smaller models fall short under retrieval: Smaller models were unable to take full advantage of the available tools when they had to retrieve them rather than being handed them.
- GPT-5 nearly closes the retrieval gap: For GPT-5, performance with tool retrieval came very close to its performance with ground-truth tools.
- Simple environments are tractable, enterprise ones are not: The most advanced reasoning models are effective at discovering tools in simpler settings but struggle seriously with complex enterprise environments.
- Scale and composition remain hard: Navigating tens of thousands of tools and combining them in non-trivial ways to solve complex problems is still a challenging task for current models, requiring both better reasoning and better retrieval models.
Methodology in Plain English
The researchers wanted a benchmark that reflects how agents would actually use real services, so they took the public REST APIs of a range of real-world services and turned them into MCP servers — collections of callable tools that an agent can invoke. This produced a very large pool of more than 18,000 tools. For every task in the benchmark, they also hand-annotated the specific tools that a correct solution needs.
They then ran agents under two conditions. In the first, the agent is simply handed the correct ground-truth tools, which isolates the question "can the agent use the tools well if it knows which ones it needs?" In the second, the agent must retrieve the right tools on its own from the full pool, which reflects the messier real-world setting. Comparing the two conditions shows how much of an agent's success depends on tool discovery versus tool use, and comparing both against browser-based agents shows whether tool calling is actually worth it. Note that the abstract does not report the specific tasks, services, model list, or any quantitative results.
Why This Matters
Research impact: The paper shifts attention from "can an agent use a tool" to "can an agent find and compose the right tools among tens of thousands," and it provides a benchmark that separates retrieval quality from reasoning quality. That separation is useful for diagnosing where future agent research should focus.
Real-world applications:
- Enterprise assistants that must operate across many internal business systems rather than through a browser.
- Customer service and CRM automation, where agents need to look up records and take actions across several services.
- Developer and IT operations tooling, where tasks require chaining multiple API calls into a coherent workflow.
- Data and analytics workflows that combine several service APIs into one multi-step job.
- Any setting where a maintained set of task-specific tools is preferable to a brittle GUI or browser automation.
Industry relevance: The paper speaks directly to the economics of agent deployment. It claims that tool-based agents can match or exceed browser-based ones and can reduce costs when tool selection is handled well — but it also warns that tool discovery at enterprise scale, not tool use, is the current bottleneck. That points to retrieval infrastructure as a first-class engineering problem alongside model capability.
Future Directions
- Better retrieval models that can operate reliably over tens of thousands of tools, since retrieval quality is the stated limiting factor for smaller models and complex environments.
- Stronger reasoning for tool composition, so agents can chain tools in non-trivial ways rather than solving only single-step or simple-environment tasks.
- Closing the retrieval-versus-ground-truth gap across all model sizes, rather than only for the most capable models such as GPT-5.
- Extending benchmark coverage of complex enterprise workflows, the setting the authors identify as where advanced models seriously struggle.
Target Audience
Researchers and engineers working on LLM agents, tool calling, and retrieval systems; benchmark designers interested in evaluation under realistic tool-selection constraints; and industry practitioners deciding whether to invest in MCP-style tool integrations versus browser-based automation for enterprise workflows.
Authors’ abstract
Since the introduction of the Model Context Protocol (MCP), the number of available tools for Large Language Models (LLMs) has increased significantly. These task-specific tool sets offer an alternative to general-purpose tools such as web browsers, while being easier to develop and maintain than GUIs. However, current general-purpose agents predominantly rely on web browsers for interacting with the environment. Here, we introduce TheMCPCompany, a benchmark for evaluating tool-calling agents on tasks that involve interacting with various real-world services. We use the REST APIs of these services to create MCP servers, which include over 18,000 tools. We also provide manually annotated ground-truth tools for each task. In our experiments, we use the ground truth tools to show the potential of tool-calling agents for both improving performance and reducing costs assuming perfect tool retrieval. Next, we explore agent performance using tool retrieval to study the real-world practicality of tool-based agents. While all models with tool retrieval perform similarly or better than browser-based agents, smaller models cannot take full advantage of the available tools through retrieval. On the other hand, GPT-5's performance with tool retrieval is very close to its performance with ground-truth tools. Overall, our work shows that the most advanced reasoning models are effective at discovering tools in simpler environments, but seriously struggle with navigating complex enterprise environments. TheMCPCompany reveals that navigating tens of thousands of tools and combining them in non-trivial ways to solve complex problems is still a challenging task for current models and requires both better reasoning and better retrieval models.