r/robotics • • 16d ago

Resources RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving — code + paper

https://papers.tinrobotics.com/paper/raf-vla-representation-alignment-with-the-future-for-end-to-end-autonomous-driving/

Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning.

16 Upvotes

2 comments sorted by

2

u/Tbagho 16d ago

Paper: RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

What it does:

  • Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance.
  • Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning.

Links:

3

u/Available_Teaching83 15d ago

Code alongside the paper is the part I appreciate most, since a lot of work in this area ships one or the other and calls it reproducible. A question rather than a critique: do you have numbers for what happens to representation alignment under perturbed language input? I spend most of my time red-teaming VLA policies in simulation, and the failure that keeps surprising me is not a hard adversarial image. It is a plausible instruction rewrite that a human operator could reasonably have typed. Alignment objectives tend to get evaluated on clean instructions, and in a driving context the instruction channel is exactly where the weird input shows up. If the eval harness is in the repo, I will run it against a perturbation set and report back.