Cheaper models widen access without removing the last mile
Price cuts, a robot-planning interface and a unified maths library make AI capability cheaper while leaving physical control and kernel optimisation as separate bottlenecks.
Artificial Intelligence··Morning
The robot planner does not drive motors directly
Google DeepMind opened Gemini Robotics ER 2 to developers through the Gemini API and Google AI Studio. The model monitors progress from video, plans multi-step tasks and coordinates multiple robots, while leaving low-level control to vision-language-action models and interfaces. Google's own measurements report 57.4% accuracy for progress classification and 91.3% for finding the critical moment; no independent verification was provided. ER 2 can also read instruments, rulers and thermometers, detect failure from video and revise its plan. Those high-level abilities do not replace lower layers governing force limits, nearby-human detection and emergency stops. For an API call to complete physical work, the plan still has to be translated into a safe controller and verified again in the real environment. Developer and production access also differ: the public API and studio are available while deployment on the enterprise agent platform remains in private preview. Model access does not mean every security, identity and fleet layer is ready.[1]
Token prices fell sharply
OpenAI cut GPT-5.6 Luna input and output prices by 80%, to $0.20 and $1.20 per million tokens respectively. Both Terra prices fell 20%, while Sol was unchanged. The company attributes the reduction to efficiency across models, inference infrastructure, production software and context management; the ratios come from the company announcement and corroborating reports, not an independent cost measurement. The price cut directly changes the budget per use, but total system cost is not only the token tariff. Long context, retries, tool calls, safety filters and latency targets determine how much compute one task consumes. Sol's unchanged price also shows that the same efficiency gain or commercial decision was not applied at every capability tier in the family. Lower unit prices shift the threshold most for products making many short calls; long tool-using workflows can rebuild the bill through call volume. The price table is an input to architecture, not total cost of ownership.[2]
Cheaper access does not remove kernel work
NVIDIA's generally available nvmath-python 1.0 package brings CUDA-X and NVPL maths libraries under one Python interface. Stateful APIs amortise planning cost across repeated calls; the company reports a 256% increase from autotuning in a particular RTX A6000 configuration. The robot layer, token price and maths kernel together show that capability can spread more widely while outcomes still depend on interface and deployment engineering. Although nvmath-python moves intensive computation into a more accessible language layer, the reported acceleration belongs to particular hardware and problem shapes. Sparse-tensor support, stateful calls and custom kernel fusion produce different gains across workloads. Wider access therefore requires three separate accounts: model-call price, numerical fit to hardware, and the engineering that connects an output to a physical or software process. Company benchmarks do not generalise without a common hardware and software base. One interface may improve portability, while memory layout, data transfer and kernel choice still shape performance. Expertise moves down a layer rather than disappearing.[3], [1], [2]