AI.info
Volker Tresp
Explore Volker Tresp on AI.info.
- ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
- AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration