AI.info
Nigel Collier
Explore Nigel Collier on AI.info.
- Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
- Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
- Confidence Estimation for LLMs in Multi-turn Interactions
- Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors