Eigen RadarAI
Analysis

AMD, Amazon and Mozilla make the case for open AI infrastructure

New FP8 software, a robot-training data loop and Mozilla’s infrastructure argument show how open tools are being positioned for cheaper, controllable AI deployment.

Artificial Intelligence··Morning
In a dark robotics workshop, a compact arm records and repeats motion over a small object inside a fine blue data-particle loop connecting varied accelerator boards.

FP8 training landings on AMD GPUs

AMD and Meta engineers have merged their FP8 training work for AMD Instinct GPUs into the main branches of TorchAO and TorchTitan, the PyTorch blog reports. On eight MI300X GPUs with Llama3-8B, rowwise FP8 delivers 13.4 per cent more throughput than BF16. The post describes kernel work that cuts data movement at three separate levels. On DeepSeek-V3 671B shapes, coalescing the writes in the colwise scales kernel brought one mixture-of-experts layer down from 7,290 microseconds to 1,170, a 6.2 times speedup. Fusing five separate forward-pass kernels into a single Triton kernel raised end-to-end throughput by 17 per cent on eight MI325X GPUs, from 5,996 tokens per second to 7,027, and recovered 89 per cent of the gap between BF16 and FP8. Adding auto-detection for AMD’s e4m3fnuz format is treated as a correctness requirement: when TorchAO scaled against the wrong maximum value, the overflow silently degraded model quality instead of raising an error. Widening the Triton autotune search space produced no measurable gain and was reverted.[1]

An open robot training data loop

Teams at Amazon and Hugging Face have combined AWS’s Apache 2.0 Strands Robots SDK, Hugging Face’s LeRobot framework and the Xet-backed storage buckets announced in March 2026 into a single workflow. A robot writes its demonstrations straight into a bucket, training streams the dataset without downloading it, and the trained weights are pushed back to the hardware. Content-defined chunking cuts the data transferred across the Hub by about four times: on a 500 MB file, changing 1 per cent of the bytes required a 5.5 MB upload. On an NVIDIA L4, 500 optimizer steps of a 51.6 million-parameter ACT policy took 133 seconds. On a 10 GB payload, a warm CDN read reached 1,086 MB per second against 780 MB per second cold. The figures come from the authors’ own runs. The loop packages capture, streaming training and weight return as open tooling rather than a closed robot stack.[2]

Mozilla frames open AI as infrastructure

Raffi Krikorian, Mozilla’s chief technology officer, told Rest of World that policymakers should look at AI as infrastructure rather than as a product. A Mozilla report published last month puts the performance gap between the best open-source models and proprietary systems at 3 per cent. Krikorian says individual users stay on ChatGPT and Claude while the workflows that IT and HR teams assemble are migrating to open models for price-performance and data-control reasons, with companies in regulated industries wanting their data to stay inside the firewall. The same report notes that Alibaba’s open-source Qwen model was downloaded more times in February 2026 than the next eight models combined. Krikorian compares the open-weight ecosystem to Linux, which began as a kernel few took seriously and ended up under Android; the debate picked up again after Meta released a new open-weight model on August 10. Set beside AMD’s FP8 kernels now in TorchAO and TorchTitan and Amazon and Hugging Face’s open robot data loop, the argument is that open software and weights are being positioned for cheaper training, controllable data paths and infrastructure-scale deployment—not only for hobby use.[3], [1], [2]

References

  1. News sourcePyTorchAMD's FP8 training work has landed in TorchAO and TorchTitan↩1↩2
  2. News sourceHugging FaceOne loop that captures robot demonstrations, trains on them and returns the weights↩1↩2
  3. News sourceRest of WorldMozilla's CTO wants policymakers to treat AI as infrastructure↩