Model support now runs from training to an Ascend package
TRL fixes loss accounting as vLLM and llama.cpp add model paths and Ultralytics adds Ascend export. The four releases expose training, runtime, and hardware stages of model support.
Artificial Intelligence··Morning
Correct accounting in post-training
Hugging Face's TRL 1.9.1 fixes incorrect normalization of DAPO, CISPO, and VESPO losses when steps per generation differ from gradient-accumulation steps. Communicator initialization in vLLM server mode is brought into line with the current API, and queue-wait measurement is corrected. The Liger path gets a fix for crashes on pre-Ampere GPUs, while DeepSpeed preparation is repaired for a CPU-offloaded optimizer. Two dependencies are temporarily pinned to avoid broken development releases.[1]
New model paths in two runtimes
vLLM 0.26 adds modeling, LoRA, speculative decoding, and NVFP4 quantization for Inkling. It can select an attention backend per KV-cache group and offload cache state to CPU memory or object storage. llama.cpp b10142 adds the MiniMax-M3 text model, vision tower, multimodal projector, and sparse-attention path. Its indexer moves to CUDA, prompt caching is enabled, and cache calculations switch to 32-bit floating point. The projects' reported speed results come from different setups and do not form a common benchmark.[2], [3]
Hardware-specific export
Ultralytics 8.4.107 exports YOLO models through CANN into static-shape FP16 `.om` packages for targets such as Ascend 310P3 or 310B4. Detection, segmentation, pose, oriented boxes, classification, and depth are covered. Export can run on Linux without an attached Ascend device; inference requires the CANN runtime and `ais_bench`. The notes give no measured speed result. Across the four updates, support extends from correct loss accounting through model runtimes to a hardware-specific package.[4]
Related columns
For more information on this topic, you can read the related columns.