Eigen RadarAI
Analysis

Open model development moves beyond one mould

Multilingual speech, alternatives to transformers, and a lower-memory distillation method show open model development diversifying across architecture, interface, and training cost.

Artificial Intelligence··Night
A violet-and-cyan sound path leaves a copper studio microphone, passes through three distinct computing structures, and reaches an irregular array of speakers.

Open weights extend into speech

NVIDIA has published its multilingual speech model Magpie TTS on Hugging Face under the company's own open model licence. It covers 12 languages: English, Spanish, French, German, Italian, Vietnamese, Mandarin, Hindi, Japanese, Modern Standard Arabic, Korean, and Brazilian Portuguese. Open weights let teams run the model on their own hardware, tune latency in their software stack, and customise it for particular uses. This form of distribution does not make the entire source code open, and use remains governed by NVIDIA's open model licence. The announcement focuses especially on the time until the first audio arrives. NVIDIA reports 32 ms to first audio on a B200 graphics processor and throughput of 320 times real time at 64 concurrent streams. The company says this interval determines whether a voice agent feels fluid or delayed. The latency and throughput figures are NVIDIA's own measurements and have not been independently reproduced. The release extends open model work into speech generation, language coverage, and live latency.[1]

Paths beyond the transformer multiply

MIT Technology Review's survey brings together startups testing different architectural routes against the rapid growth in computing cost as text becomes longer. Miami-based Subquadratic is working on SubQ, which uses sparse attention instead of comparing every word with every other word. Manifest AI is testing an approach that keeps rolling summaries as a conversation develops instead of carrying the entire context. Liquid AI combines transformers with its liquid neural networks and reports 34 million downloads of its models. Inception follows a diffusion-based route that generates text in blocks rather than in sequence; the company claims Mercury 2 produces results comparable with some of OpenAI's GPT-4 models while running 10 times faster. Pathway replaces the attention mechanism with state-space mathematics. These startups do not converge on one new architecture. The survey instead presents separate responses to the same computing-cost problem: sparse attention, rolling summaries, liquid networks, diffusion, and state spaces. Their performance claims come from the companies and have no independent evaluation in the report. Model development is widening through several computational methods for long context and text generation.[2]

Distillation fits into a smaller training setup

A method published by Multiverse Computing targets two bottlenecks in shrinking an existing model family with fewer resources. The first change computes and caches the teacher model's top 100 logits once, so the teacher does not remain in memory throughout training. The second processes data in chunks instead of materialising at once a loss matrix as large as vocabulary size multiplied by sequence length. According to the team's measurements, at 32K tokens peak memory falls from 85.2 to 5.45 gibibytes, step time for GPT-OSS 20B drops from 57.0 seconds to 12.23 seconds, and the training setup fits on 1 GPU node instead of 4. The implementation is available as open-source code, while the method is described in arXiv:2608.03796, a preprint that has not been peer reviewed. A student with 3.2 billion parameters, distilled from a teacher with 8 billion, is said to retain most of the teacher's accuracy on BoolQ and HellaSwag, though the page provides no figures for either dataset. Magpie TTS offers a downloadable speech interface, the transformer alternatives diversify architecture, and this work describes a smaller training setup. Openness appears at three distinct levels across the reports: distributing weights, rebuilding architecture, and sharing a training method in code.[1], [2], [3]

References

  1. News sourceHugging FaceNVIDIA publishes Magpie TTS, a 12-language speech model, with open weights↩1↩2
  2. News sourceMIT Technology ReviewA cluster of startups is testing architectures meant to replace the transformer↩1↩2
  3. News sourceHugging FaceA distillation recipe cuts peak training memory from 85.2 to 5.45 gibibytes↩