Skip to content
AI.info

Research

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Overview Research area: Evaluation of LLM-based web-browsing agents; multilingual and multimodal information retrieval and benchmark design. Technical level: Intermediate (accessible to readers famili

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
arXiv
2610.03574
Published
2026-10-02
Authors
Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina

AI summary

Overview

  • Research area: Evaluation of LLM-based web-browsing agents; multilingual and multimodal information retrieval and benchmark design.
  • Technical level: Intermediate (accessible to readers familiar with LLM agents and search benchmarks; no specialized mathematics required).
  • Scope in one sentence: The paper introduces HyperBrowseComp, a 423-question, 13-language, multimodal benchmark designed to stress-test how well web-browsing agents can find and connect obscure, publicly verifiable evidence on the open web.

What This Paper Is About

Existing web-browsing benchmarks tend to isolate one challenge at a time: the original BrowseComp is English-centric and mostly text-based, language extensions such as BrowseComp-ZH and K-BrowseComp cover single non-English web environments, and XBCP uses translated supporting documents while keeping questions and answers in English. HyperBrowseComp's goal is to combine natively authored multilingual questions with evidence that spans many modalities — webpages, images, video, audio, scanned documents, maps, tables, and figures — so that agents must first figure out what and where to search, not just retrieve an answer from a well-known source.

Key Contributions

  1. A new benchmark: HyperBrowseComp, containing 423 questions manually authored by native or highly proficient speakers of the corresponding languages, spanning 13 languages, and validated by a second annotator with multi-model difficulty auditing.
  2. A multilingual, multimodal, open-web design: Unlike translation-based or fixed-corpus benchmarks, questions are authored directly in each language and grounded in that language's web environment, with evidence drawn from heterogeneous sources including video, scanned documents, images, maps, and audio.
  3. A harness comparison: Evaluation of 5 web-search-enabled frontier LLMs (Gemini 3.1 Pro Preview, Gemini 3.7 Flash, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.6 Sol) across three search integration methods (provider built-in search, Exa, and OWL), analyzing accuracy, token cost, wall-clock time, and failure cases.
  4. A human baseline: A human evaluation on 30 sampled questions (10 each in Indonesian, Thai, and Vietnamese) to contextualize model accuracy and effort.

Main Findings

  • Headroom is large. The strongest configuration, Gemini 3.7 Flash with built-in search, achieved 31.68% accuracy (134 correct), followed by Gemini 3.1 Pro Preview at 25.53% (108 correct). GPT-5.6 Sol, Terra, and Luna with built-in search scored 19.15%, 15.84%, and 11.58% respectively.
  • Most questions defeat all five native-search runs. 244 of 423 questions (57.68%) had no recorded correct answer across the five models with built-in internet search; among the 179 questions answered by at least one model, 49 were solved by exactly one.
  • Search harness choice matters as much as the model. Under Exa, GPT-5.6 Sol improved to 26.71% (113 correct) over its built-in 19.15%, whereas both Gemini variants dropped sharply: Gemini 3.7 Flash fell to 22.22% (94 correct) and Gemini 3.1 Pro to 18.20% (77 correct).
  • OWL was hampered by runtime instability. OWL with Gemini 3.7 Flash scored 17.73% (75 correct); 93 terminal workforce failures were counted as incorrect. On completed attempts alone, Gemini 3.7 Flash reached 22.73% with OWL, still below its built-in result.
  • OWL's workload was heavy but cheaper in tokens. OWL with Gemini 3.7 Flash made 11,486 external tool calls (27.15 per attempt): 6,953 DuckDuckGo searches, 3,719 visual-browser calls, 375 Wikipedia searches, 206 PDF queries, 141 video queries, 47 image queries, 30 image-to-text calls, and 15 file reads. It made 56,342 model calls (56,398 attempts) with 28 model-level errors. The 93 terminal failures broke down as 51 worker-step timeouts, 28 exhausting the 250-call model budget, nine worker processing failures, four hitting the 2,400-second wall-clock limit, and one other runtime failure.
  • Cost does not track accuracy. The three Exa configurations used 1.32–2.20 million tokens per question, whereas OWL used about 462 thousand tokens and 133 model calls per attempt. Gemini's internal search harness was both better-performing and cheaper than Exa or OWL.
  • Humans and models perform comparably on a 30-question sample. Across all 30 attempts (including give-ups), human mean elapsed time was 89.5 minutes, median 54.6 minutes; accuracy was 15/30. Participants submitted answers for 26 questions and gave up on four. Gemini 3.7 Flash scored 13/30 on the same subset, and both answered correctly on only 8 questions together.
  • Human difficulty does not mirror model difficulty. Humans solved Indonesian questions accurately and quickly (6/10, mean 40.23 min, median 34.68 min) while Gemini scored 3/10; humans struggled on Thai (3/10, mean 159.70 min, median 60.33 min) and Gemini also scored 3/10; Vietnamese was easiest for Gemini (7/10) but took humans longer (6/10, mean 68.56 min, median 57.92 min).
  • Difficulty concentrates in visual and specific domains. Shared failures clustered on visual modalities such as plots, graphs, and images, and on domains including Games and Food & Drink, with the most frequent model failures on German and Indonesian queries. A TF-IDF comparison of English translations associated terms such as "magazine," "page," and "photograph" with shared failures; the authors describe these patterns as descriptive, not causal.
  • Dataset composition. Approximately 64.3% of questions were classified as requiring at least one non-text modality. The most frequent source categories were Video/YouTube (39.0%), PDF/book/OCR (29.8%), and Image (18.4%), with an Arithmetic modality tag on 28.1% of questions (categories overlap). Domains were led by Economics & Business (18.9%), Media & Entertainment (15.8%), Education (11.3%), and Geography & Transport (9.0%); the section text describes 12 primary domains plus an Other category (the conclusion refers to 13 primary domains), with five questions (1.2%) classified as Other.
  • Benchmark construction filtered out 77 questions. A no-internet audit using seven models (Gemini 3.1 Pro Preview, Gemini 3.7 Flash, GLM-4.7, GPT-OSS-120B, Kimi K2 Thinking, MiniMax-M2, DeepSeek-V3.2) found 121 questions answered correctly by at least one model, including 54 answered by at least two. The authors excluded 54 questions answered by two or more models plus 23 more answered by exactly one model but judged to require highly specific answers.

Methodology in Plain English

The authors started from the answer rather than the question. For each candidate item, an author identified a publicly verifiable fact and its supporting evidence, then wrote a question that requires nontrivial searching to recover that fact. Authors submitted the question, the canonical answer, acceptable aliases, relevant URLs, exact evidence locations (a quotation, page number, table or figure identifier, video timestamp, audio interval, or map location), source language, relevant modalities, access dates, and a solution trace.

A second annotator validated each item, checking interpretability, answer correctness and uniqueness, evidence accessibility, and time-sensitive wording. Validators could request non-leading "fingerprints" — constraints intended to disambiguate rather than to help search, such as "the book contains exactly 13 illustrations of cats" — and could ask for the difficulty to be raised.

To keep questions from being answerable from parametric knowledge alone, the team ran seven models with no internet access and removed questions that any model could answer, using a combined rule: remove if two or more models got it right, or if exactly one did and the answer required a highly specific value. That removed 77 questions, leaving 423.

Models were then evaluated under a shared agent protocol. Built-in search and Exa used a standard ReAct loop capped at 25 research steps, where the model drives its own search queries and browsing. Exa provided standardized web_search and web_fetch tools (up to 5 results per query, page content truncated to 20,000 characters, 60-second timeouts). OWL ran as a separate multi-agent workforce with task manager, coordinator, web specialist, multimodal specialist, and synthesis specialist roles, using DuckDuckGo, Wikipedia, a headless Chromium visual browser, and image/video/PDF/audio inspection tools. Answers were judged for binary correctness by an LLM judge given the question, full response, and reference answer; the authors report that using each model as its own judge introduced at most 0.80 percentage points of aggregate variance across judge sets.

The dataset is released in encoded form with a decoding key to reduce accidental ingestion by automated data collection pipelines.

Why This Matters

Impact on research. HyperBrowseComp argues that web-agent performance cannot be reported as a single model number — the retrieval harness is a first-class variable, since swapping built-in search for Exa helped one model and hurt two others by large margins. It also provides a deliberately difficult, multilingual, multimodal target that is designed to remain discriminative longer than trivia-style QA benchmarks, and it pairs model scores with token and time costs so efficiency can be weighed alongside accuracy.

Real-world applications:

  • Building search assistants that must reconcile evidence across languages and formats, such as scanned reports, video timestamps, and map data.
  • Auditing the cost and reliability of agent harnesses before deployment, given that OWL's 93 terminal failures out of 423 attempts reflect operational fragility, not just reasoning ability.
  • Guiding provider and tool decisions — the paper shows proprietary models are tightly co-adapted to their native search ecosystems, which matters when teams substitute a third-party retrieval stack.
  • Supporting human-in-the-loop research workflows where a person and an agent have complementary strengths, as suggested by the low overlap (8 of 30) between human and model correct answers.

Industry relevance. Search providers, agent framework builders, and enterprises deploying browsing agents can use the harness comparison to reason about where accuracy gains come from — the model, the retriever, or the orchestration layer — and about whether a cheaper harness might outperform a more expensive one.

Future Directions

  • Standardizing difficulty across languages. The authors note that intrinsic question difficulty may vary across language splits, since German and Indonesian queries see the most model failures, and call standardizing difficulty an important consideration for future iterations.
  • Designing efficient, accurate harnesses. The finding that Gemini's cheaper internal search outperformed Exa and OWL motivates further work on browsing harness design as an orthogonal research direction.
  • Broader model and harness coverage. The authors state that computational cost (a single run can require days of execution and hundreds of millions of tokens) and restrictive API access limited the study to five models and only a subset of models under Exa; they propose comparing search tools while holding the model fixed under comparable resource budgets.
  • Separating sources of failure. The paper notes its descriptive patterns do not isolate whether failures come from retrieval, source access, interpretation, or reasoning, leaving that decomposition open.
  • Understanding individual human search skill. The authors flag that annotators may search differently from one another as an interesting orthogonal question outside the scope of this work.

Target Audience

Researchers and engineers working on LLM agents, web search, tool use, and multilingual or multimodal retrieval; benchmark designers interested in difficulty auditing and contamination control; and product teams evaluating search integrations and agent harnesses. Readers who want a conceptual, non-mathematical account of how a hard browsing benchmark is built, validated, and stress-tested will also find it accessible.

Authors’ abstract

We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.

Read the original paper