The local AI stack is shrinking from three directions
Muse Glimmer, ExecuTorch, and a new knowledge-distillation method show three separate ways to fit models, runtime, and training onto more limited hardware.
Artificial Intelligence··Evening
The model and runtime meet at the same target
Meta released Muse Glimmer, a model with 30 billion parameters, as an open-weight system that handles images and text together. Its weights are available through Hugging Face, and the Apache 2.0 licence permits commercial use. Meta says it designed the model for always-on local agent workflows and made it small enough to run on a Mac or PC with one consumer GPU. That positioning now has a concrete runtime path: the ExecuTorch team released ready-to-run bundles for NVIDIA CUDA and Apple Metal. In PyTorch's own measurement on an M5 Pro with 64 GiB of memory, the model generated 21.6 tokens per second; DFlash speculative decoding raised that rate to 33.0. In the model's long-context structure, 13 of its 52 layers attend to the full sequence while 39 use a sliding window. The first release supports text and image input, while video, cross-session prefix sharing, checkpointing, and continuous batching remain outside its current scope. Meta's model comparisons and PyTorch's speed results are vendor measurements, with no independent evaluation yet.[1], [2]
On the training side, the teacher leaves memory
A method published by Multiverse Computing on Hugging Face changes two expensive steps in transferring knowledge from a large teacher model to a smaller student. First, the teacher's top 100 logits are computed once and cached. In the published method, the teacher model is not kept in GPU memory throughout training. The KL loss is then calculated in chunks instead of materialising a matrix as large as vocabulary size multiplied by sequence length. In the team's own 32K-token test, peak memory fell from 85.2 GiB to 5.45 GiB, while the step time for GPT-OSS 20B declined from 57.0 seconds to 12.23 seconds. The team reports that training in the same method ran on one GPU node rather than four. The implementation is open source, and the method is described in arXiv:2608.03796, a preprint that has not been peer reviewed. The accuracy claim is narrower: a student with 3.2 billion parameters distilled from a teacher with 8 billion parameters is said to retain most of the teacher's performance on BoolQ and HellaSwag, but the announcement gives no figures for either evaluation.[3]
Efficiency does not sit behind one switch
The three releases pursue more accessible AI at different stages. Muse Glimmer shapes model size and attention for local operation; ExecuTorch carries that model onto specific hardware backends through prepared bundles; the distillation method reduces the memory burden of both the teacher and the loss calculation while training a smaller model. The result is a resource chain rather than one performance number: storage and memory footprint, token generation speed, long-context cost, and the number of GPUs needed for training are treated separately. The reports also make the limits visible. The Muse Glimmer and ExecuTorch results come from Meta and PyTorch, and the first runtime release lacks several planned capabilities. The distillation work quantifies large reductions in memory and step time, but provides no detailed figures for retained accuracy and has not completed peer review. Even with those qualifications, the published weights, runnable bundles, and open-source implementation give developers three separate entry points for testing the claims on their own hardware. Local deployment here is a combination of model design, runtime support, and a cheaper path to producing the smaller models that applications can use.[1], [2], [3]