Eigen RadarAI
Analysis

AI deployment is being measured in runtime, task cost and memory tiers

Amap’s 24-hour inference claim, DeepSeek’s per-test cost and new PCIe 6.0 storage expose three different system measures for deploying artificial intelligence.

Artificial Intelligence··Midday
Synthetic computing scene where a glass flow circles a ceramic ring, separates into measured packets and enters a three-layer porous memory structure

Amap says it brought duration to a consumer card

Amap, Alibaba’s mapping unit, says its ABot-World-0 interactive world model sustained 24 hours of uninterrupted inference on one consumer-grade graphics card. The company says mainstream interactive world models stop after roughly a minute and attributes the difference to a training method called LongForcing. Keeping the weights open through Hugging Face and Reactor gives developers a way to download the model. The announcement does not identify the card, however, or specify scene resolution, frame rate, memory use or the interactive workload maintained across the 24 hours. Both the duration and the comparison with one minute are Amap’s measurements, and no independent reproduction accompanied them. Those omissions make it harder to test which operating conditions support the claim. The news therefore presents a concrete product claim about bringing long-running interactive inference closer to consumer hardware, while leaving out the configuration another team would need to recreate the same operating conditions.[1]

DeepSeek is being priced by the completed task

Artificial Analysis estimates that DeepSeek V4-Flash completes benchmark tests at an average cost of 3 cents. On the same test set it puts Claude Fable 5 at $3.15, GPT-5.6 Sol at $1.86 and Kimi K3 at 86 cents. The ranking does not rely only on the unit price displayed by a provider. It also counts the input tokens processed and output tokens generated to finish each task. V4-Flash lists at $0.14 per million input tokens and $0.28 per million output tokens. Because the model was released on 31 July, the figures describe a research firm’s estimate from its first days in use. This measure differs from Amap’s runtime claim: it compares the token expense of a completed test, while the capital cost of hardware and the location of memory in the wider system do not enter that quoted price directly.[2]

The memory tier adds a separate hardware calculation

An assessment published before FMS considers faster enterprise SSDs as a layer between an accelerator’s expensive high-bandwidth memory and cold storage. Kioxia’s CM10 series uses a sixth-generation PCIe interface, offers direct cold-plate liquid cooling alongside air cooling and, on company figures, reaches peak sequential reads of 28.4 GB/s. Price and general availability have not been announced; drives are sampling to selected customers. PCIe 6.0 reaches 64 GT/s per lane, while forward error correction adds about 4-8 nanoseconds of latency in each direction. The assessment estimates that a 70 billion parameter model with a one-million-token context creates about 320 gigabytes of key-value cache for one user, four times the stated high-bandwidth memory capacity of an H100. Read together, the three developments show why deployment cannot be reduced to one model price. The duration of a running workload, the tokens consumed to complete a task and the movement of data between memory tiers create separate resource calculations, each attached to a different part of the operating system around the model.[3], [1], [2]

References

  1. News sourcePR NewswireAmap says its interactive world model ran for 24 hours on one consumer GPU↩1↩2
  2. News sourceKhaleej TimesA research firm measures DeepSeek V4-Flash as the cheapest well-known model per test↩1↩2
  3. News sourceTech TimesA liquid-cooled enterprise SSD moves to the centre of the AI memory-tier debate↩