Liquid AI, Modular, and PyTorch push inference beyond one hardware path
Liquid AI’s draft models, the open-source Mojo compiler, and IBM Spyre adapters seek to reduce inference cost through decoding speed, accelerator portability, and model compatibility.
Artificial Intelligence··Midday
Draft models make the same family run faster
Liquid AI released three DSpark draft checkpoints for the LFM2.5 family. In the company’s measurements, speculative decoding raises throughput by as much as 3.18 times on a GPU and 2.87 times on device. On an H100, the model with 2.6 billion parameters averaged a 2.67-times speedup, moving from 323 to 864 tokens per second, while average function-calling latency for the same model fell by 57 percent. The small draft heads ship in Safetensors and GGUF formats with day-one support for llama.cpp and SGLang. All performance figures come from Liquid AI’s own measurements.[1]
Mojo opened as the accelerator list widened
At ModCon 2026, Modular made the Mojo programming language and compiler fully open source under Apache 2.0 and moved its Modular Cloud inference service into general availability. The MAX layer above Mojo remains source-available, while the platform now supports Nvidia, AMD, AWS Trainium, Google TPUs, Apple Silicon, and Qualcomm datacenter chips. The company is aiming to open one software path across different accelerators. Forbes notes that the reported AMD MI355X comparison with Nvidia B200 also comes from Modular’s own measurements, so the verified development is the portability scope while the performance advantage remains a company claim.[2]
Thirteen adapters carried thousands of models onto Spyre
According to the PyTorch blog, 13 adapters written with AI agents made 7,960 of HuggingFace’s 10,000 most-downloaded embedding models runnable on IBM’s Spyre accelerator. In measurements from mid-April through late June, 6,804 of the models with an adapter passed the end-to-end test. The adapters construct equivalent operations that the Spyre compiler can handle without changing what the model computes. The blog also says compiler fusion makes it difficult to locate where an error begins and that agents can draw overly confident conclusions from a single experiment; broad compatibility therefore does not remove the need for human-supervised debugging.[3]