A tier taken off sale is a capacity reading
OpenAI stopped taking new subscriptions for its 200 dollar a month Pro plan. The reason product lead Thibault Sottiaux gave is load: demand for Astra, released on 3 September, is straining the infrastructure, and the Pro tier places the heaviest load on the systems of any tier. Access through the application programming interface and the cheaper Go and Plus plans stayed open, and existing Pro subscribers keep their access.[1]
Pausing new subscriptions at the most expensive tier while keeping the cheaper ones open points at a serving limit; a pricing decision does not take this shape. The company is choosing to leave revenue it already has customers for. The quantity deserves care too: what gets measured here is the number of concurrent requests that can be served today, rather than installed hardware in total or announced capacity.[1]
Two ways to get more answers out of the same racks
Amazon SageMaker Inference added prefix-aware routing, which sends requests beginning with the same text to the same instance and reuses the cached key-value pairs for that prefix. In measurements AWS published for Llama 3.1 70B on 7 ml.p5.48xlarge instances, median time to first token on long-context work fell by 71 per cent to 77 per cent, the cache hit rate rose from roughly 25 per cent to 82 per cent, and throughput on the same workload rose 15 per cent to 16 per cent. The boundary is stated and the figures are the provider's own measurements.[2]
Model caching takes the second route. On Amazon SageMaker HyperPod the model weights are downloaded to a node's local NVMe disk in advance and the inference server container image is pulled before a pod asks for it. AWS reports about 60 per cent faster scale-out with the weights cache, reads at roughly 7 GB/s, and the removal of a download that otherwise runs past 30 minutes for a model the size of DeepSeek-R1, above 600 GB. What this buys back is the time an instance spends doing no work, rather than raw compute.[3]
These three developments work on one quantity: the instance time spent per delivered answer. Prefix-aware routing takes a second computation of the same prefix out of that denominator and model caching takes out the wait for a download, while the tier OpenAI paused marks the point where the same denominator no longer divides into what is available. Another explanation is possible: the Pro pause could be commercial prioritisation, since interface customers and subscription customers need not share one hardware pool.[1], [2], [3]
What does the cost per output already show?
OpenDesign Arena scored 13 models on web apps, dashboards, mobile screens and landing pages. GPT-6 Astra came first with 82.7 points at 1.61 dollars and 11.1 minutes per design, while DeepSeek V4.1 Flash scored 81.2 points at 0.023 dollars and 5.3 minutes per design. The gap in points is 1.5; the gap in minutes is more than twofold.[4]
Those amounts are the list prices the ranking paid, so they measure what a buyer pays, and the provider's own cost stays out of sight. The minutes column says something different: that time is how long an instance is actually held for one output, independent of pricing policy. The tighter a provider's capacity, the more that column does the work the price list cannot.[4]
A trackable threshold follows. OpenAI gave no date for reopening Pro subscriptions; a reopening of sales by 31 December 2026 would show that the squeezed quantity is concurrent serving capacity and that this capacity can be grown. No reopening by the same date would point at a limit sitting deeper than compute allocation, on the memory or bandwidth side. The observable event is a single one: whether new Pro subscriptions reopen by 31 December 2026.[1]