Skip to content
AI.info

Research

Found-RL: foundation model-enhanced reinforcement learning for autonomous driving

Overview Research area: Reinforcement learning for end-to-end autonomous driving, combined with foundation models (specifically Vision-Language Models). Technical level: Advanced. The paper assumes fa

Found-RL: foundation model-enhanced reinforcement learning for autonomous driving
arXiv
2602.10458
Published
2026-02-11
Authors
Yansong Qu, Zihao Sheng, Zilin Huang, Jiancong Chen, Yuhao Luo, Tianyi Wang, Yiheng Feng, Samuel Labi, Sikai Chen

AI summary

Overview

Research area: Reinforcement learning for end-to-end autonomous driving, combined with foundation models (specifically Vision-Language Models).

Technical level: Advanced. The paper assumes familiarity with reinforcement learning training loops, reward shaping, policy optimization, and Vision-Language Models.

Scope (one sentence): The paper presents Found-RL, a platform that makes it practical to use foundation-model supervision inside reinforcement learning for autonomous driving by decoupling slow VLM reasoning from the fast simulation loop and distilling VLM guidance into a lightweight, real-time policy.

What This Paper Is About

Reinforcement learning is a leading approach to end-to-end autonomous driving, but it learns inefficiently from experience and produces policies whose decisions are hard to interpret in complex driving scenarios. Vision-Language Models carry rich, context-aware knowledge that could compensate for these weaknesses, but they are too slow to sit inside the tight, high-frequency training loops RL depends on. Found-RL's goal is to capture the benefit of VLM knowledge without paying its latency cost during learning.

Key Contributions

  1. An asynchronous batch inference framework that decouples heavy VLM reasoning from the simulation loop, removing the inference-latency bottleneck and enabling real-time learning rather than blocking training on per-step VLM calls.
  2. Two supervision mechanisms for distilling VLM guidance into the RL policy: Value-Margin Regularization (VMR) and Advantage-Weighted Action Guidance (AWAG), which transfer expert-like VLM action suggestions into the learned policy.
  3. Dense reward shaping using high-throughput CLIP, with Conditional Contrastive Action Alignment to address CLIP's "dynamic blindness." This conditions prompts on discretized speed and command, producing a normalized, margin-based bonus derived from context-specific action-anchor scoring.
  4. A complete end-to-end pipeline for integrating fine-tuned VLMs into RL, released alongside code, data, and models at the project's public repository.

Main Findings

  • Latency can be decoupled from learning: The asynchronous batch inference design is presented as the core innovation that resolves the VLM latency bottleneck for high-frequency RL training, supporting real-time learning.
  • VLM knowledge can be distilled into a compact policy: The abstract reports that a lightweight RL model achieves near-VLM performance compared with billion-parameter VLMs, suggesting large models are not required at deployment time.
  • Real-time inference is sustained: The lightweight model runs at approximately 500 FPS according to the abstract, indicating suitability for the kind of high-frequency operation driving requires.
  • CLIP needs contextual grounding for reward shaping: The paper identifies a limitation it calls dynamic blindness in CLIP-based rewards and proposes Conditional Contrastive Action Alignment, which conditions on discretized speed and command to produce a normalized, margin-based bonus instead of a raw similarity score.
  • Details beyond these claims are not in the abstract: No specific benchmark scores, dataset sizes, ablation results, or comparisons against named baselines are reported in the abstract, so the magnitude of the performance claim cannot be assessed from it alone.

Methodology in Plain English

The researchers built a training platform rather than a single model. Instead of asking a large Vision-Language Model to comment on every simulation step — which would stall training — they let the VLM run separately in batches and feed its judgments back asynchronously. Those judgments are then converted into learning signals in two ways: a regularization term that nudges the policy's value estimates using the margin implied by VLM-preferred actions (VMR), and an advantage-weighted scheme that steers the policy toward VLM-suggested actions in proportion to how much better they appear to be (AWAG). A separate, much faster vision-language model, CLIP, is used to generate dense rewards throughout training. Because plain CLIP similarity does not account for the fact that the right action depends on the current driving situation, they condition the CLIP prompts on discretized speed and command, then score candidate action "anchors" for that context and turn the result into a normalized, margin-based bonus. Finally, they provide the plumbing to fine-tune and plug in a VLM end to end, so the resulting RL policy stays small and fast enough for real-time use.

Why This Matters

Impact on research: The work reframes the foundation-model-for-RL question from "can it help?" to "how do you make it fast enough to be useful?" Its asynchronous inference pattern and its treatment of CLIP's context blindness are reusable ideas for any RL setting where a large model provides supervision but cannot run at control frequency.

Real-world applications:

  • Autonomous driving, where low-latency decision-making is a hard requirement and large-model reasoning is normally confined to offline or low-frequency components.
  • Robotics and embodied control, where the same latency mismatch between foundation models and control loops appears.
  • Simulation-based training pipelines that need dense, semantically meaningful reward signals instead of hand-crafted or purely geometric rewards.
  • Deployment on compute-constrained vehicles, since the distilled lightweight policy can run without a billion-parameter model in the loop.

Industry relevance: The result that a lightweight policy can approach VLM-level behavior at real-time speed is directly relevant to automotive and robotics companies weighing whether foundation-model knowledge can appear in shipped systems, not just in research prototypes. The public release of code, data, and models lowers the barrier to reproducing and adapting the pipeline.

Future Directions

  • Verifying the "near-VLM performance" claim across benchmarks, scenarios, and baselines that the abstract does not specify — including how much of the gap to a full VLM remains and in which conditions.
  • Testing whether the asynchronous batch inference schedule degrades when VLM feedback lags behind rapidly changing scenarios, and how stale guidance interacts with policy updates.
  • Extending the Conditional Contrastive Action Alignment approach to additional conditioning variables beyond discretized speed and command, and assessing how sensitive the reward bonus is to the discretization scheme.
  • Evaluating transfer from simulation to real vehicles, since the abstract describes an end-to-end simulation-oriented pipeline and does not report real-world deployment.

Target Audience

Researchers and graduate students working on reinforcement learning, autonomous driving, or foundation-model integration into decision-making systems. It is also relevant to practitioners in autonomous-vehicle and robotics engineering who need to understand the latency trade-offs involved in putting large models anywhere near a control loop. Readers without a background in RL training or Vision-Language Models will find the paper demanding, since the contributions assume familiarity with both.

Authors’ abstract

Reinforcement Learning (RL) has emerged as a dominant paradigm for end-to-end autonomous driving (AD). However, RL suffers from sample inefficiency and a lack of semantic interpretability in complex scenarios. Foundation Models, particularly Vision-Language Models (VLMs), can mitigate this by offering rich, context-aware knowledge, yet their high inference latency hinders deployment in high-frequency RL training loops. To bridge this gap, we present Found-RL, a platform tailored to efficiently enhance RL for AD using foundation models. A core innovation is the asynchronous batch inference framework, which decouples heavy VLM reasoning from the simulation loop, effectively resolving latency bottlenecks to support real-time learning. We introduce diverse supervision mechanisms: Value-Margin Regularization (VMR) and Advantage-Weighted Action Guidance (AWAG) to effectively distill expert-like VLM action suggestions into the RL policy. Additionally, we adopt high-throughput CLIP for dense reward shaping. We address CLIP's dynamic blindness via Conditional Contrastive Action Alignment, which conditions prompts on discretized speed/command and yields a normalized, margin-based bonus from context-specific action-anchor scoring. Found-RL provides an end-to-end pipeline for fine-tuned VLM integration and shows that a lightweight RL model can achieve near-VLM performance compared with billion-parameter VLMs while sustaining real-time inference (approx. 500 FPS). Code, data, and models will be publicly available at https://github.com/ys-qu/found-rl.

Read the original paper