Eigen RadarAI
Analysis

Nvidia puts a 30-fold per-megawatt claim at the centre of its inference push

Nvidia claims up to 30 times more work per megawatt for Vera Rubin, Cerebras speeds up the same chip, and Groq 3 LPX enters production with Nebius.

Artificial Intelligence··Evening
Light streams spread from one luminous server rack across a dark data center, symbolizing inference density and energy efficiency.

A 30-fold per-megawatt claim for Vera Rubin

Nvidia says Vera Rubin NVL72 can deliver up to 30 times more agentic-AI inference work per megawatt than GB300 NVL72. The company measured the result with AgentX, an open-source benchmark in SemiAnalysis's InferenceX suite. AgentX replays prerecorded coding sessions and captures long-context prefill and the gaps between tool calls. The same Nvidia post says GB300 NVL72 delivers up to 15 times more throughput per megawatt than H200 NVL8 on DeepSeek V4-Pro while cutting cost per million tokens by up to 10 times. These are Nvidia's own measurements, and the company says the Vera Rubin comparison is awaiting review by SemiAnalysis.[1]

Cerebras gives the same chip more power

Cerebras keeps the previous generation's 5-nanometre WSE-3 chip in its CS-4 system. The company raised the clock rate and added power and cooling instead of changing the chip. One rack now holds three wafers instead of two. The company reports up to 4,400 tokens per second per user and says the system is up to 30 times faster than setups using Nvidia GPUs. Those comparisons are Cerebras's own measurements. SemiAnalysis analysts who reviewed the claims found the networking gains small. Memory remains 44 gigabytes per wafer, and the report says OpenAI uses Cerebras hardware for Codex Spark.[2]

Groq 3 LPX reaches full production

Nvidia's Groq 3 LPX inference accelerator, announced at Hot Chips 2026, has entered full production, with cloud provider Nebius Group N.V. as its first production customer. Nvidia positions it as an extension of the Vera Rubin data-centre platform and supports up to 256 accelerators in a full rack. Artificial Analysis measured 3,400 tokens per second on Gemma 4 31B with a 100,000-token context window. Nvidia licensed Groq Inc.'s technology for 20 billion dollars in December 2025 and hired founder Jonathan Ross and president Sunny Madra. The company has not disclosed pricing or a general-availability date.[3]

References

  1. News sourceNVIDIA Developer BlogVera Rubin promises 30 times the work per megawatt, on Nvidia's own runs↩
  2. News sourceTHE DECODERCerebras doubled its speed without changing the chip↩
  3. News sourceSiliconANGLEGroq 3 LPX enters full production with Nebius as first buyer↩