An agent model on one card and an open speech model
Meta Superintelligence Labs says it published Muse Glimmer, a 30 billion parameter model, under the Apache 2.0 licence. The company says the model is for always-on agents that run on a Mac or PC without a cloud connection; weights sit on Hugging Face, with llama.cpp, MLX and ExecuTorch builds described as arriving in the coming days. According to Meta, Glimmer was distilled from the larger Muse Spark model through longer-context, agent-heavy mid-training and post-training that combines supervised fine-tuning, on-policy distillation and reinforcement learning. Full precision would need more than 55 gigabytes of memory; roughly 4 bit quantisation brings it under 20 gigabytes, which Meta says fits cards with 24 gigabytes to 32 gigabytes with minimal to no degradation on agentic tasks. The model takes interleaved text and images, was trained on data from more than 100 languages, and works with scaffolds including OpenClaw. Meta reports strong results against Gemma4-31B and Qwen3.6-27B on DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench; those figures are the company's evaluation and are not independently verified. In the same window, NVIDIA's Hugging Face post introduced Magpie text-to-speech with open weights, a NVIDIA NIM package and 12 languages, adding Modern Standard Arabic, Korean and Brazilian Portuguese. The company says it targets teams that manage voice-agent latency inside their own infrastructure.[1], [2]
