Qwen-Drive 1.0 cuts road veering, while its explanations can miss the manoeuvre
Alibaba's Qwen team has released Qwen-Drive 1.0, a model with 4 billion parameters that combines traffic-scene understanding and vehicle-motion planning in one framework. Reinforcement learning cut its simulated road-veering rate from 24 percent to 12 percent. However, the reason it gives does not always match the manoeuvre it plans, and scenarios requiring different reactions can still be confused. Code, weights and demo data are available under Apache 2.0.
Artificial Intelligence··Evening
Perception and planning in one model
Alibaba's Qwen team and Huazhong University of Science and Technology have released Qwen-Drive 1.0-4B to the research community. The system retains the Qwen3.5-4B base and adds two components: a bird's-eye-view perception head for three-dimensional object detection, occupancy and road segmentation, and a planning expert that generates the vehicle's trajectory. The Decoder and TechNode report the same release as an attempt to combine traffic-scene understanding and vehicle-motion planning within one framework.[1], [2]
Road veering falls while the reasoning gap remains
In the simulation results reported by The Decoder, reinforcement learning reduced the road-veering rate from 24 percent to 12 percent. The model outperformed the unmodified Qwen3.5-4B on driving question answering and three-dimensional perception, with limited loss on general-knowledge tasks. Its stated explanation still does not always match the manoeuvre it subsequently plans, and it can confuse situations that require very different responses. These figures are a model comparison within the reported simulation evaluation, not a real-road safety outcome.[1]
Two planning routes and an open release package
TechNode distinguishes two planning versions: one learns by imitating driving examples, while the other adds reinforcement learning to that training. The release provides code, model weights and demo data together under the Apache 2.0 licence. The Decoder also reports that the weights are on Hugging Face, the code on GitHub and the accompanying paper on arXiv. Researchers can therefore inspect the architecture and reported results, though the availability of those files does not itself resolve the mismatch between explanations and manoeuvres.[2], [1]