Skip to content
AI.info
News
Tools
Research
Learn
AI.info
Angelika Romanou
Explore Angelika Romanou on AI.info.
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
Page 1