Skip to content
AI.info

Research

Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models

Overview Research area: Natural Language Processing — multilingual evaluation, predictive performance estimation, and agentic LLM systems. Technical level: Intermediate. Familiarity with LLM evaluatio

Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
arXiv
2604.08970
Published
2026-04-10
Authors
Avni Mittal, Shanu Kumar, Sandipan Dandapat, Monojit Choudhury

AI summary

Overview

Research area: Natural Language Processing — multilingual evaluation, predictive performance estimation, and agentic LLM systems.

Technical level: Intermediate. Familiarity with LLM evaluation metrics (BLEU, chrF, pass@1, MAE) and basic agent/retrieval concepts helps, but the paper explains its setting clearly enough for a non-specialist reader.

Scope: The paper introduces a controlled 1,500-question benchmark for estimating missing Task–Model–Language evaluation results, plus a DAG-orchestrated agentic system, Litmus (Re)Agent, that predicts those results from restricted literature evidence.

What This Paper Is About

Multilingual evaluation coverage is sparse: for many task–model–language combinations there is no published benchmark score, or the available evidence comes from other languages, other models, or inconsistent reporting conditions. The authors frame this as the Task–Model–Language (TML) prediction problem and ask how well a model will perform on a task in a target language when direct evaluation results are missing. They build a benchmark that deliberately hides part of the evidence and a retrieval-and-reasoning agent system that must infer the missing result from what remains.

Key Contributions

  1. A controlled benchmark for predictive multilingual evaluation. 1,500 questions spanning six tasks (code generation, mathematical reasoning, QA/VQA, text classification including NLI, text summarisation, machine translation) and five evidence scenarios (S1–S5), with 50 questions per task–scenario block (25 PredSet, 25 QnASet). The benchmark separates accessible evidence (a reduced paper corpus available at inference) from ground truth (defined from a larger combined corpus).

  2. Two complementary query types. PredSet requires predicting a scalar score for a task–model–language query; QnASet requires comparative reasoning to identify the best-performing model or language.

  3. Litmus (Re)Agent, a DAG-orchestrated, citation-grounded agentic system built on the earlier LITMUS architecture, with three targeted improvements: enriched expert knowledge (a broader multi-task cross-lingual corpus), enhanced code execution via lang2vec and the URIEL typological database, and prompt stabilisation for expert-aligned reasoning.

  4. A comprehensive empirical study: comparison against five baselines, per-task and per-scenario breakdowns, internal agent-behaviour diagnostics, an LLM-judge quality analysis, a backbone-LLM ablation, and a within-subjects human evaluation with 8 participants.

Main Findings

  • Best overall system: Litmus (Re)Agent achieves the lowest overall PredSet MAE at 10.4 in Table 2 (the text of Section 5.1 reports 10.3), ahead of LITMUS++ (12.7), Single Agent (14.5), ThoughtAgent (15.1), GPT-4.1 (16.5 in Table 2; 16.4 in the text), and Magentic-One (17.2). On QnASet it reaches 21.6% accuracy, versus 20.1% for Single Agent, 15.9% for both ThoughtAgent and GPT-4.1, 15.1% for LITMUS++, and 10.0% for Magentic-One. (Section 5.1 describes the Single Agent QnASet figure as 18.9%, which does not match Table 2.)

  • Largest gains where evidence is weakest: In scenario S4 (distant language, same model) Litmus (Re)Agent holds MAE at 11.7 while Magentic-One degrades to 23.5; in S5 (different language and model) it reaches 11.9 versus 15.7 for LITMUS++.

  • Code generation shows the widest margin: PredSet MAE 14.2 for Litmus (Re)Agent versus 39.8 for GPT-4.1, and QnASet accuracy 55.3% versus 27.2% for GPT-4.1.

  • Mathematical reasoning is the exception: It is the only task where the GPT-4.1 direct baseline performs best on PredSet MAE (18.2), which the authors attribute to weaker dependence on external evidence.

  • Multi-agent chat alone is not enough: The Single Agent (MAE 14.5) outperforms the ThoughtAgent group chat (15.1), suggesting that multi-agent coordination without DAG-level decomposition does not always help. In QnASet scenario S2, the Single Agent reaches 22.0% while the ThoughtAgent drops to 9.3%.

  • Scenarios are not a strict difficulty ladder: PredSet MAE is lowest in S3 (9.0) rather than S1 (9.4), indicating that close-language transfer can provide cleaner signals than noisy direct evidence.

  • Internal reasoning quality is uneven: Across 1,730 conversations and 5,184 generated hypotheses, Litmus (Re)Agent scores 93.9% on thought faithfulness, 91.6% on feature correctness, 83.5% on web search relevance, and 70.0% on capability compliance — but only 30.6% on code execution success. Allowing retries raises execution success to 54.6%.

  • LLM-judge quality profile: Litmus (Re)Agent leads on predictive plausibility (4.6) and feature selection (4.8) on a 1–5 scale, while GPT-4.1 leads on coherence (5.0) and citation emphasis (4.3). LITMUS++ scores markedly lower across all dimensions, which the authors read as evidence for the impact of the three improvements.

  • Metric type predicts difficulty: Accuracy-based metrics show the highest MAE for Litmus (Re)Agent (12.7), compared with ROUGE (6.5), chrF (7.5), and BLEU (5.0).

  • Backbone matters but is not the whole story: In a code-generation ablation across S1, S3, and S5, GPT-4.1 gives the best quality scores, o4-mini is closest, DeepSeek-V3 is strong in S3 but drops in S1 and S5, and Llama-3.3-70B performs worst (e.g., citation emphasis 1.1–1.3). The authors conclude gains come mainly from the structured, evidence-grounded workflow rather than the backbone alone.

  • Humans rate the system's output higher: In the 8-participant study, the largest improvements over general LLM tools were in justification (+1.0) and actionability (+0.9). Novices gained more overall (confidence +1.0, actionability +1.0), while experts showed the largest gain in justification (+1.1). Novices took longer with the system (13.7 vs. 5.9 min per question); expert time was stable (10.3 vs. 10.7 min). In the final survey, 5 of 7 participants reported improved prediction accuracy.

  • Reported inconsistencies: Several figures in the narrative text do not match the tables (overall PredSet MAE 10.3 vs. 10.4; GPT-4.1 MAE 16.4 vs. 16.5; Single Agent QnASet 18.9% vs. 20.1%; text summarisation described as MAE 4.1–7.8 with Single Agent best at 4.1, while Table 2 lists 4.7–8.3 and Table 6 gives a Single Agent average of 7.1 for that task).

Methodology in Plain English

Building the benchmark. The authors collected multilingual evaluation papers and extracted language-to-model-family mappings, aggregating them into a task-specific "combined" mapping that represents the full answer space. They then created a "reduced" mapping by removing a subset of papers. Systems only ever see the reduced corpus; the ground-truth answers come from the combined set. This is what makes the benchmark "controlled" — the experimenters know exactly what information is missing and how much.

Defining evidence scarcity. Five scenarios vary what is observable. S1 has the target language and target model family present (direct evidence). S2 keeps the language but removes the model. S3 and S4 remove the language but keep the model, with S3 allowing transfer from a typologically close language and S4 only from a distant one. S5 removes both. "Same model" means the same model family, not the exact checkpoint. Language similarity for S3 and S4 is computed as cosine distance over lang2vec features, split into close and distant groups by a percentile-based threshold.

Normalising scores. Ground-truth values were automatically extracted from the paper collection and manually validated, then normalised to a 0–100 scale: values in [0, 1] are multiplied by 100, and other metrics undergo task-specific linear scaling. The recorded metrics include pass@1, Accuracy, F1, and BLEU, and evaluation is restricted to compatible metric families within a task.

How the system works. A query such as "How will GPT-4o perform on machine translation in Nepali?" goes to a MainAgent, which decides whether it is new or a follow-up. New queries go to a ThoughtCreatorAgent, which breaks the problem into hypotheses and spawns one ThoughtAgent per hypothesis as a node in a directed acyclic graph. Each ThoughtAgent coordinates five specialised sub-agents: a Research Planner, Web Search and Crawl, Expert Knowledge (a curated multilingual knowledge base), a Coder that writes and executes analysis scripts including regression models, and a Reporter. A ThoughtAnalyzerAgent monitors progress, spawns new hypotheses when gaps appear, and prunes redundant branches (agents move between Active, Completed, and Discarded states). Finally, a ResponseAnalyzerAgent aggregates everything into a prediction with citations and a structured rationale. Each hypothesis keeps an independent evidence trail, which limits context-window saturation.

The three improvements over the predecessor. The expert knowledge base was expanded from a small set of task-specific papers to a broader corpus of multi-task cross-lingual evaluation studies encoding language–task–model observations, failure modes, and analysis patterns. The Coder agent gained lang2vec and URIEL, providing syntactic, phonological, geographic, and family-level features for over 7,000 languages, enabling language distance computation and feature-informed regression. Prompts were redesigned to enforce expert-style workflows and cited evidence, reducing invalid tool calls, brittle code execution, and hallucinated citations.

Evaluation setup. All systems run under the same reduced-corpus restriction; the restriction is part of the protocol, not a system limitation. Six systems are compared: Litmus (Re)Agent, LITMUS++ (the prior DAG system without the improvements), ThoughtAgent (a single multi-agent group chat with no DAG decomposition), Single Agent (a ReAct loop with direct tool access), a GPT-4.1 direct baseline with no agentic scaffolding, and Magentic-One, a general-purpose multi-agent framework. All agentic systems use GPT-4.1 as the backbone, served through the Azure OpenAI API (version 2025-01-01-preview), with temperature 1.0 and top-p 1.0; the GPT-4.1 direct baseline uses temperature 0.7 and a 2,000-token output limit. The Expert KB agent runs at temperature 0, filters retrieved documents by cosine similarity with a threshold of 0.90, and selects the top k = 2 results. Tooling includes Autogen for orchestration, ChromaDB as the vector store, and the Firecrawl API for search and scraping.

Why This Matters

Impact on research. The paper reframes multilingual evaluation as a prediction problem under explicitly controlled evidence conditions, rather than only a coverage problem. By separating accessible evidence from ground truth, it gives the community a reusable testbed for measuring how well systems reason from incomplete scientific literature — and its negative findings (group chat without DAG decomposition does not help; code execution is a bottleneck at 30.6%) point to concrete engineering targets. It also connects prior feature-based transfer prediction work (LangRank, Task2Vec, LEEP, LogME) with literature-grounded agentic reasoning.

Real-world applications:

  • Model selection for multilingual deployment: practitioners can get a reasoned estimate of expected performance before committing to expensive evaluation runs.
  • Localisation and market prioritisation: teams can assess which languages a candidate model is likely to handle well, and where coverage gaps are worth funding.
  • Low-resource language planning: the benchmark explicitly targets settings where published evidence is thin, using typologically close and distant language transfer as signals.
  • Decision-support with citations: every prediction comes with supporting citations and a reasoning trace, which matters for review and audit.

Industry relevance. The author affiliations (Microsoft, MBZUAI, IIT Hyderabad) reflect direct industry interest in multilingual product deployment. The workflow is directly applicable to teams that must choose among models for dozens of languages without budget to benchmark each combination, and the paper's comparisons — including a general-purpose multi-agent framework (Magentic-One) that performs poorly without domain-specific design — give a realistic picture of how much task-specific engineering is still required. The ethical framing is also relevant to practice: the authors explicitly state predictions are decision-support signals, not substitutes for benchmarking when high-stakes decisions require ground-truth measurement, and warn that estimates can inherit biases from uneven language coverage and inconsistent reporting in the published literature.

Future Directions

  • Fixing the code execution bottleneck. Success at the conversation level is 30.6%, rising to only 54.6% with retries, even though hypothesis generation and evidence grounding are reliable (93.9% thought faithfulness). Better code-generation and self-repair strategies are an obvious next step.
  • Testing beyond GPT-4.1. The authors note that generalisation to smaller or open-source backbones is unverified, and their own ablation shows meaningful variation across GPT-4.1, o4-mini, DeepSeek-V3, and Llama-3.3-70B-Instruct.
  • Broadening benchmark scope. Coverage spans six tasks but excludes multimodal reasoning, dialogue, and safety, and depends on a fixed curated corpus; expanding both the task space and the evidence sources remains open.
  • Disentangling architectural from prompt-level gains. The comparison against LITMUS++, Magentic-One, and the ThoughtAgent uses different internal coordination strategies, so reported gains mix architecture and prompting effects.
  • Improving metric-aware calibration. Since accuracy-based metrics show substantially higher MAE (12.7) than ROUGE (6.5), BLEU (5.0), and chrF (7.5), the authors suggest metric-aware normalisation or task-specific calibration as a route to better predictions.

Target Audience

Researchers and engineers working on multilingual NLP evaluation, low-resource language technology, and LLM agent architectures. It will be especially useful to practitioners who must make model-selection or deployment decisions in languages where direct benchmark results do not exist, and to benchmark designers interested in controlled protocols for studying reasoning under deliberately restricted evidence. Readers focused on agentic systems will find the DAG orchestration design, the internal agent diagnostics, and the negative results about group chat coordination particularly relevant.

Authors’ abstract

We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are missing. This problem is common in multilingual deployment, where evaluation coverage is sparse and published evidence is uneven across languages, tasks, and model families. We introduce a controlled benchmark of 1,500 questions spanning six tasks and five evidence scenarios. The benchmark separates accessible evidence from ground truth, enabling evaluation of systems that must infer missing results from incomplete literature evidence. We also present Litmus (Re)Agent, a DAG-orchestrated agentic system that decomposes queries into hypotheses, retrieves evidence, and synthesises predictions through feature-aware aggregation. Across six systems, Litmus (Re)Agent achieves the best overall performance, with the largest gains in transfer-heavy scenarios where direct evidence is weak or absent. These results show that structured agentic reasoning is a promising approach to multilingual performance estimation under incomplete evidence.

Read the original paper