Project · Active
VLAlert: vision-language models for adaptive driver alerting
What it is
VLAlert (VLAlert-Bench) is a vision-language framework and a unified per-tick benchmark for adaptive driver alerting: each second, a model looks at the last 8 frames of video and decides whether to stay silent, observe, or alert. The benchmark integrates six driving-event datasets — Nexar Collision, DoTA, DAD, DADA-2000, ADAS-TO-Critic, and a Kaggle accident set — into 13,534 videos and 192,892 labeled one-second ticks.
Why it matters
Most driver-assistance systems treat alerting as binary and reactive: an alert fires once risk crosses a threshold, with no notion of building attention gradually or holding back when a driver is already responding appropriately. VLAlert’s three-state framing — and pooling six disparate event datasets into one labeling scheme — makes it possible to train and evaluate alerting policies that escalate the way an attentive human passenger would, rather than defaulting to blunt binary warnings.
Connections
VLAlert applies the alerting problem that TriDrive’s driver-forecasting work targets from the world-model side, and its underlying event sources overlap with ADAS-TO. It sits in the lab’s AI for mobility and Vehicle technologies threads.
Project team
- Hao Zhou (PI)
- Yuhang Wang
- Lingyao Li (University of Arizona)
Outputs & releases
- Preprint: Wang, Zhou. Bridging human oversight and black-box driver assistance: vision-language models for predictive alerting in lane keeping assist systems. arXiv:2505.11535 — see also the publication page
- Project page: wangyuhang-cmd.github.io/vlalert
- Dataset: huggingface.co/datasets/HenryYHW/VLAlert
- Venue: Conference on Robot Learning (CoRL) 2026
- License: CC-BY-4.0