Autonomous driving · Challenge 2026
AlpaSim E2E Closed-Loop Challenge 2026
A closed-loop benchmark for evaluating autonomous driving policies across safety, generalization, and computational efficiency.
Organized by HKU OpenDriveLab, NVIDIA ASPIRE Group, and KE:SAI.

From trajectory matching to closed-loop driving
How do we know whether an autonomous driving model is truly getting better? Matching a recorded trajectory is only part of the answer. A useful evaluation must also ask whether a vehicle avoids collisions, stays on the road, completes its task, and produces decisions within the computing limits of a real vehicle.
AlpaSim brings these questions into a shared, open evaluation framework. Every predicted trajectory changes the simulated vehicle’s state. The simulator then renders new sensor observations and feeds them back to the policy, allowing evaluation to follow the consequences of successive decisions.
- 01
Observe the environment
- 02
Predict a trajectory
- 03
Update the vehicle state
- 04
Render new observations
This continuous loop tests whether a policy can recover from deviations, whether errors accumulate over time, and how a mistaken decision affects the rest of a drive.
Why move beyond open-loop evaluation and NAVSIM?
A recorded drive is one valid solution, not the only safe way to navigate a road. Safely slowing down can produce a larger Average Displacement Error (ADE) than an unsafe trajectory that ends closer to the recording. Trajectory error alone can therefore reward the wrong behavior.

NAVSIM evaluates safety, comfort, and driving progress through lightweight simulation. Since 2024, it has supported three public challenges; the first attracted 143 teams and 463 submissions. NAVSIM v2 added pseudo-simulation, using 3D Gaussian Splatting to create observations at different positions, headings, and speeds and assess recovery beyond the recorded trajectory.
AlpaSim takes the next step: a full sensor-level loop in which a policy continuously observes, acts, and receives fresh observations generated from the state it has reached.
Two complementary tracks
The challenge pairs evaluation at industry scale with an accessible path for reproducible academic research. A common developer kit supports development and training on both datasets.
Industry-scale evaluation
Physical AI AV Track
Built on the NVIDIA Physical AI Autonomous Vehicles Dataset, with approximately 1,700 hours of driving data across diverse regions and environments. NuRec supports closed-loop testing of stability, generalization, and computational efficiency at scale.
Reproducible research
nuPlan Track
Built on the nuPlan ecosystem, extending the evaluation direction of WorldEngine with MTGS reconstruction assets. Researchers can adapt NAVSIM-style models for method comparisons, model iteration, and closed-loop behavior analysis.




Evaluation that accounts for deployment
16 GiB
GPU memory limit
≤ 0.1 s
Target model work per Drive call
Docker
Containerized policy submissions
The organizers provide the computing resources and runtime environment for official evaluation. Each team receives the same number of official submission opportunities, while standardized local validation sets support debugging and ablation studies. Observation and control frequencies differ between tracks; all submissions must meet the challenge’s runtime requirements.
Beyond an average score
The challenge is exploring Item Response Theory (IRT) to estimate policy ability and scenario difficulty from patterns of success and failure. Treating policies as test takers and scenarios as questions helps reveal differences in difficulty and discrimination, and provides uncertainty intervals alongside rankings.
Physical AI AV and nuPlan are scored independently. For policies with overlapping rank intervals, the challenge further considers the average distance driven per at-fault infraction.
Challenge timeline
June 15, 2026
Challenge opens
Start developing and evaluating your driving policy.
September 15, 2026
Planned rules freeze
Rules and submission formats freeze following maintenance.
October 31, 2026
Final submissions
Public leaderboard closes; final submissions and technical reports are due.
Dates follow the published challenge plan. See the official challenge website for the latest schedule, rules, and submission details.
Make progress measurable
Join the community in developing safer, more efficient, and more generalizable end-to-end driving policies—and help make every improvement reliably measurable, reproducible, and verifiable.