Research
BRIDGE: Predicting Human Task Completion Time From Model Performance
Overview Research area: AI evaluation and psychometrics — specifically Item Response Theory (IRT) applied to large language model benchmarking, and forecasting of AI capability in human-interpretable

- arXiv
- 2602.07267
- Published
- 2026-02-06
- Authors
- Fengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau, Siva Reddy, Hugo Larochelle
AI summary
Overview
Research area: AI evaluation and psychometrics — specifically Item Response Theory (IRT) applied to large language model benchmarking, and forecasting of AI capability in human-interpretable units.
Technical level: Intermediate. The paper assumes familiarity with IRT (the two-parameter logistic model), log-odds reasoning, and benchmark evaluation harnesses, though the core argument is accessible to anyone who understands benchmark scores and human task-time estimates.
Scope: The paper introduces BRIDGE, a framework that fits a 2PL IRT model to model–task success/failure data across several benchmark suites, calibrates the resulting latent task-difficulty scale against human completion times, and uses that calibration to predict human task durations for new benchmarks and to forecast frontier model task-length horizons without new human studies.
What This Paper Is About
Benchmark scores are hard to interpret in terms of real-world effort: a score gain might reflect improvements on short, routine tasks while longer multi-step tasks remain out of reach. METR's approach of measuring AI capability as the length of tasks a system can complete with 50% probability (the "50%-task-completion time horizon") solves this interpretability problem but depends on expensive, noisy human time annotations, and requires a bespoke human study for each benchmark. BRIDGE asks whether human task completion time can instead be predicted from model performance data alone, by learning a latent difficulty scale from model responses and anchoring it to human completion time.
Key Contributions
-
A unified psychometric framework: BRIDGE fits a two-parameter logistic (2PL) IRT model to model–task binary outcomes across multiple benchmarks, jointly estimating each task's discrimination parameter ('a'), each task's latent difficulty ('b'), and each model's latent ability ('θ').
-
An empirical anchoring of latent difficulty to human time: The paper shows that latent task difficulty varies approximately linearly with the logarithm of human completion time (fit R² = 0.81 on METR tasks), where a one-unit increase in 'b' corresponds to about a 2.26× increase in human completion time.
-
Out-of-distribution time prediction without new human studies: Using the calibration learned on METR, BRIDGE estimates human completion times on SWE-bench Verified, MLE-bench, GDPval, and Cybench, and validates these against available annotations and qualitative expectations.
-
Capability forecasting from model performance alone: BRIDGE forecasts frontier model task-length horizons in human time units without human time annotations, reproducing METR's exponential scaling result with a doubling time of approximately 6 months.
Main Findings
-
Latent difficulty tracks log human time: Across the METR task suites (SWAA, HCAST, RE-Bench), the fitted log-linear relationship between IRT difficulty 'b' and log human completion time has R² = 0.81, with each one-unit increase in 'b' corresponding to roughly a 2.26× increase in human completion time.
-
BRIDGE beats heuristic and LLM baselines on SWE-bench Verified: On the annotated coarse time buckets (<15 minutes, 15 minutes to 1 hour, 1 hour to 4 hours, >4 hours), BRIDGE produces monotonically increasing predictions that track bucket ordering and scale. The logit success-rate baseline systematically underestimates task duration, while Gemini 3 Pro and GPT-5.2 produce compressed predictions with limited dynamic range and poor separation between longer-horizon buckets.
-
Strong alignment on Cybench: BRIDGE achieves the strongest correlation with recorded human first-solve times (R² = 0.45) and places 92.3% of tasks within a 0.5×–2× tolerance band. The logit success-rate baseline substantially underestimates duration; the LLM estimators (Gemini 3 Pro, GPT-5.2) capture qualitative ordering but consistently overestimate absolute times.
-
MLE-bench difficulty ordering: Tasks yielding only a valid submission are estimated to require substantially shorter human completion times than achieving above-median leaderboard performance or earning any Kaggle medal.
-
Frontier model horizons at 50% success: Models released in 2025 achieve 50% success on tasks estimated to require roughly 1–2.5 hours of human effort. The current state-of-the-art model reaches a task-length horizon of roughly two hours at a 50% success rate.
-
Raising the threshold contracts the horizon: At an 80% success threshold, even the most recent models are largely limited to tasks requiring less than one hour of human completion time, though the doubling time is similar.
-
Exponential growth with a ~6-month doubling time: Partitioning models into 2-month release windows and selecting the highest-ability model in each, BRIDGE finds that the solvable human task-length horizon at 50% success grows exponentially, with capabilities doubling approximately every 6 months — corroborating METR's human-time-based findings using only model performance data.
-
Robustness to sparse response matrices: In a sparsity ablation removing 10%–70% of observed model–task entries, the correlation of task difficulty estimates with full-data estimates stays above 0.97 after removing 50% of observations and above 0.94 at 70% removal (0.9424 at 70% sparsity, 15% observed). Model ability estimates are even more stable (0.9798 at 70% sparsity). The paper notes the actual experiments use a response matrix that is 51% observed.
-
Benchmark-dependent discrimination: Curves are markedly steeper for METR and SWE-bench than for MLE-bench and GDPval, consistent with higher discrimination parameters. Non-smoothness in probability–time curves arises from heterogeneity in task-level difficulty and discrimination within a benchmark.
-
Disagreement with Epoch AI on frontier model ranking: Epoch AI's IRT-based framework (L2-regularized non-linear least squares, anchor at the model level) ranks Gemini 3 Pro as the top-performing model, whereas BRIDGE's task-level estimates place Claude 4.5 Opus ahead. The authors hypothesize this stems from different benchmark compositions.
Methodology in Plain English
The researchers gathered binary success/failure records for many model–task pairs across several benchmark suites. For METR's aggregated dataset of 170 tasks spanning SWAA, HCAST, and RE-Bench, repeated trials per pair were collapsed into a single binary outcome, with a pair counted as successful if the model succeeded in at least 50% of reported trials.
They then fit a 2PL IRT model to this response matrix. IRT, originally developed for educational testing, assumes each task has a difficulty and a discrimination parameter and each model has an ability parameter, and predicts the probability of success as a logistic function of the gap between ability and difficulty. The model was fit by maximizing likelihood using Markov chain Monte Carlo with hierarchical priors. Since IRT parameters are only identifiable up to a scale, the resulting difficulty values have no inherent meaning on their own.
The key step is calibration. For the subset of tasks with human completion-time annotations (METR), the researchers regress the logarithm of human completion time on latent difficulty 'b' to obtain a log-linear mapping. This resolves the scale ambiguity and gives the latent axis a human-interpretable unit. They then apply that same mapping to tasks in new benchmarks with no time annotations, producing estimated human completion times from model performance alone.
For forecasting, they take the best-performing model in each release window, use its estimated ability as the difficulty of the hardest task it can solve at a 50% success rate, map that difficulty through the calibration to a human completion time, and plot the resulting horizon over model release dates.
Validation includes comparing against coarse annotated time buckets on SWE-bench Verified, against human first-solve times on Cybench, and against two baselines: a logit success-rate heuristic (which transforms task and model average success rates into log-odds scores) and prompted frontier LLMs (Gemini 3 Pro and GPT-5.2) asked to estimate expert human completion time from task descriptions, using greedy decoding with a 32,000 maximum output token budget.
Evaluation used the InspectAI framework with a default ReAct-style scaffold, bash shell and Python interpreter tool access, generation capped at 1000 turns or until the context window filled, temperature 0.0, and sufficiently large maximum output token budgets. Cybench runs used the Cybench agent framework in the unguided setting, restricted to tasks for which at least one model succeeds. GDPval used an LLM-as-a-judge pipeline with Gemini 3 Pro as the judge because the benchmark does not admit a fully automated verifiable evaluator.
Why This Matters
Impact on research: The paper suggests that a latent difficulty scale estimated purely from model performance can substitute for direct human time annotations, lowering the cost and friction of tracking real-world AI progress. It provides independent validation of METR's exponential task-horizon scaling result using an entirely different data source, and it opens a psychometric route to human-interpretable evaluation that can transfer across heterogeneous benchmarks.
Real-world applications (as implied by the framing):
-
Capability tracking and forecasting: Organizations can monitor how the length of tasks AI can reliably complete at a given success threshold changes over model releases, expressed in hours of human effort rather than opaque benchmark points.
-
Benchmark selection and interpretation: Teams choosing which benchmarks to run can convert raw scores into estimated human task durations, making score changes comparable across different benchmark formats and domains.
-
Agent deployment scoping: The 50% and 80% success-threshold analyses give a rough sense of which task durations an agent can handle reliably, informing where humans still need to remain in the loop.
-
Evaluation cost reduction: Because new benchmarks can be mapped onto the calibrated time scale without new human studies, evaluation pipelines can scale to newly introduced benchmarks faster.
Industry relevance: The framework targets the practical bottleneck of grounding benchmark numbers in real-world effort. The contrast with Epoch AI — where the two methods rank frontier models differently (Gemini 3 Pro versus Claude 4.5 Opus) — underscores that capability forecasting carries inherent uncertainty and that multiple independent methodologies are useful for calibrating confidence.
Future Directions
-
Human–AI collaborative settings: Extending BRIDGE to model how task responsibility is dynamically shared, how partial automation reshapes effective task difficulty, and how human interventions can serve as signals of difficulty as model capabilities advance.
-
Knowledge-intensive evaluations: Enriching the psychometric model with task attributes capturing information requirements, external tool use, or reliance on knowledge beyond a model's training cutoff, moving beyond long-horizon procedural tasks.
-
Uncertainty-aware difficulty estimation: As task horizons lengthen, human completion times are expected to show increased inter-individual variability; accounting for this variability is a natural extension toward uncertainty-aware forecasting.
-
Reconciling divergent forecasts: Understanding why BRIDGE and Epoch AI produce different frontier-model rankings, and whether the differences trace to benchmark composition or modeling choices, would sharpen the reliability of both approaches.
Target Audience
Researchers and practitioners in AI evaluation and benchmarking, particularly those interested in psychometric methods, agent capability measurement, and capability forecasting. The paper is also relevant to policy and safety analysts who track long-horizon AI progress, and to benchmark designers who want their scores expressed in human-interpretable units. Readers without a background in IRT will need to work through the log-odds formulation, but the downstream results and forecasting methodology are stated in accessible terms.
Authors’ abstract
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Existing approaches that rely on direct human task completion time annotations are costly, noisy, and difficult to scale across benchmarks. In this work, we propose BRIDGE, a unified psychometric framework that learns a latent difficulty scale from model responses and anchors it to human task completion time. Using a two-parameter logistic Item Response Theory model, we jointly estimate latent task difficulty and model capability from model performance data across multiple benchmarks. We demonstrate that latent task difficulty varies linearly with the logarithm of human completion time, allowing human task completion time to be inferred for new benchmarks from model performance alone. Leveraging this alignment, we forecast frontier model capabilities in terms of human task length and independently reproduce METR's exponential scaling results, with the 50% solvable task horizon doubling approximately every 6 months.