China's Tsinghua End-to-End Autonomous Driving Model iDriveVLA Ranks First Globally in NAVSIM
2026-07-21 10:12
Favorite

en.Wedoany.com Reported - The iDriveVLA, a one-stage multimodal end-to-end autonomous driving model independently developed by the team of Academician Li Keqiang from Tsinghua University's School of Vehicle and Mobility (led by Professor Li Shengbo and Assistant Researcher Guan Yang), has achieved the world's top ranking in the international NAVSIM autonomous driving benchmark. The model scored 94.85 on the core comprehensive metric PDMS, marking significant progress in end-to-end foundation models following the team's development of China's first full-stack end-to-end autonomous driving system.

NAVSIM Challenge Leaderboard

iDriveVLA is a one-stage multimodal end-to-end model that addresses the challenge of balancing trajectory completeness and selection accuracy in autonomous driving models through two key innovations. First, it constructs a hierarchically decoupled cross-attention neural network model architecture, comprising atomic components such as an image encoder, multimodal trajectory decoder, and evaluation selection decoder. These components are optimized through a combination of macro-level driving semantics and micro-level performance scores, coupled with a probabilistic gating filter mechanism for low-performance trajectories, enhancing the model's cognitive understanding of driving safety and compliance requirements. Second, the training platform employs a two-stage ladder training method proposed by the team, consisting of supervised pre-training followed by reinforcement learning fine-tuning: initially, the model's basic driving capability is established through imitation learning of expert driving trajectories; subsequently, three reinforcement learning key technologies—harmonious gradient BPO for resisting multi-source data conflicts, exponentially weighted hybrid loss function Hybrid-Loss, and the STAPO fine-tuning algorithm for eliminating spurious tokens—are utilized to achieve high-completeness generation and high-accuracy evaluation of multimodal trajectories in urban traffic scenarios. This achievement reflects the team's sustained innovation capability in vehicle-road-cloud integrated autonomous driving foundation models, providing technical support for developing highly safe, reliable, and generalizable autonomous driving systems.

This research was supported by the Key Program of the National Natural Science Foundation of China, the Regional Joint Project of the Beijing Natural Science Foundation, and Dongfeng Motor Corporation.

The NAVSIM Challenge was co-initiated by NVIDIA, Stanford University, the University of Tübingen, and the Shanghai Artificial Intelligence Laboratory, and was officially announced at CVPR 2024, a top-tier conference in the field of artificial intelligence. It serves as an important public benchmark in end-to-end autonomous driving. According to incomplete statistics, it has attracted over 143 teams and 463 models in recent years, bringing together renowned domestic and international enterprises such as NVIDIA, Bosch, Valeo.ai, Horizon Robotics, Xiaomi Auto, Changan Automobile, Huawei Yinwang, and Qianli Technology, as well as prestigious universities including UCLA, the University of Toronto, EPFL, the University of Hong Kong, and Fudan University. The challenge uses the Predictive Driver Model Score (PDMS) as its core technical metric, evaluating the autonomous driving capability of submitted models by comprehensively calculating multidimensional indicators such as collision risk, driving progress, compliance, and comfort within a fixed time horizon.

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com