Eigen RadarAI
Analysis

Open weights and distillation push capable models toward one-card memory

Meta's 30 billion parameter Muse Glimmer, NVIDIA's 12-language Magpie speech model and Multiverse's distillation memory method place side by side the drive to fit open-weight systems onto narrower hardware.

Artificial Intelligence··Midday
In a bright workshop, fine model pathways and a voice waveform compress toward one graphics card in an open workstation, with cables and tools on the wooden bench.

An agent model on one card and an open speech model

Meta Superintelligence Labs says it published Muse Glimmer, a 30 billion parameter model, under the Apache 2.0 licence. The company says the model is for always-on agents that run on a Mac or PC without a cloud connection; weights sit on Hugging Face, with llama.cpp, MLX and ExecuTorch builds described as arriving in the coming days. According to Meta, Glimmer was distilled from the larger Muse Spark model through longer-context, agent-heavy mid-training and post-training that combines supervised fine-tuning, on-policy distillation and reinforcement learning. Full precision would need more than 55 gigabytes of memory; roughly 4 bit quantisation brings it under 20 gigabytes, which Meta says fits cards with 24 gigabytes to 32 gigabytes with minimal to no degradation on agentic tasks. The model takes interleaved text and images, was trained on data from more than 100 languages, and works with scaffolds including OpenClaw. Meta reports strong results against Gemma4-31B and Qwen3.6-27B on DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench; those figures are the company's evaluation and are not independently verified. In the same window, NVIDIA's Hugging Face post introduced Magpie text-to-speech with open weights, a NVIDIA NIM package and 12 languages, adding Modern Standard Arabic, Korean and Brazilian Portuguese. The company says it targets teams that manage voice-agent latency inside their own infrastructure.[1], [2]

Distillation's memory ceiling and layered speech stacks

In a Hugging Face post, Multiverse Computing describes two changes that lower distillation's memory peak when shrinking large language models. The usual setup keeps teacher and student models in memory together and reruns the teacher at every step. With gpt-oss-120b and its 201,088 token vocabulary, at sequence length 32K and batch size 4, the teacher probability tensor alone takes about 50 gigabytes in bfloat16, and one training iteration can peak at roughly 250 gigabytes. That peak exceeds the 141 gigabytes on an H200 card. The first change caches the teacher's 100 most likely tokens per position once; the second folds the output projection into the loss so the student's full vocabulary grid is never built. The team says the second method brings the peak to about 128 gigabytes and makes growth linear with sequence length, while computing the output projection twice. The measurements are the post's own. On Magpie, NVIDIA says integrated speech models are simpler because one call carries audio in and out, while a cascaded stack of speech recognition, text-to-speech and language model components keeps each layer tunable. It recommends that structure for data residency, latency visibility and domain adaptation; the post is a product introduction; performance comparisons are the company's own.[3], [2]

Open weights and tight memory on the same table

The three developments sit on different product layers, yet centre the same hardware squeeze. Muse Glimmer says quantisation brings memory need into the 24–32 gigabyte range of one consumer card and shares weights under an open licence; that claim rests on the company's evaluation. The Multiverse post argues regenerating teacher output at every distillation step can exceed H200 memory, and that its measurements bring the peak to about 128 gigabytes. Magpie offers an open-weight speech model across 12 languages for teams managing latency in their own infrastructure, and ties integrated versus cascaded setups to the need for control. The shared line is fitting the model into a local or on-premises memory budget and distributing weights or components openly. The reader gets not a ranking, but the limits companies report when parameter count, quantisation, caching and loss accounting shrink the memory ceiling. Independent verification sits outside these texts.[1], [3], [2]

References

  1. News sourceMeta AI ResearchMeta opens the weights of Muse Glimmer, a 30 billion parameter agent model that runs on one consumer graphics card↩1↩2
  2. News sourceHugging FaceNVIDIA publishes its Magpie speech model with open weights and support for 12 languages↩1↩2↩3
  3. News sourceHugging FaceMultiverse Computing publishes a method that brings distillation's memory peak within reach of one graphics card↩1↩2