AI.info
Guanhua Chen
Explore Guanhua Chen on AI.info.
- Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence
- Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
- BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
- From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
- Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
- ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
- G2: Guided Generation for Enhanced Output Diversity in LLMs