The frontier model race
We follow the performance race among the most advanced artificial intelligence models from major technology companies. This covers leaps in reasoning capabilities and new features of frontier models.
Briefs
- DeepSeek approaches US models on LiveBench, with a reported gap of 3 percent
- Six model routers fail to beat random selection in an experiment
- BitNest reduces memory by putting its draft inside shared weights
- Altman defends broad AI access despite accepting some risks
- Vercel confirms KVM flaw reported through its Sandbox program
- Strata adds parallel conversations and a local connection for Codex
- Research finds over 100 rural US data-center projects could qualify for tax breaks
- Robinson calls for safety-culture changes after leaving OpenAI
- Attack attempts reported after disclosure of Mythos-linked HFS flaw
- Muse instructions call for profiles of people around each user
- Governments endorse AI-driven research in the Kyoto Vision
- Free Gemini accounts lose Flash and Pro access on 9 October
- DoorDash details how shared gateways manage AI models and agents
- Claude asks for separate permission to train on voice recordings
- Claude Code 2.1.289 fixes plugin and shell permission bypasses
- Buterin tests private AI advice and finds limits in speed and data sharing
- Sam Altman calls surrendering human judgment to AI a safety issue
- Aleph Alpha opens Kolibri weights for German and English
- Hans Anders suspends Meta glasses sales amid privacy debate
- Anthropic’s disclosure tally separates Claude findings from shipped fixes
- ChatGPT opens connected-account finance questions to US Free and Go users
- Bonsai World turns satellite images into robot training environments
- Amazon releases Strands Decider for choices inside AI workflows
- NVIDIA prepares easier model launching across local AI systems
- Microsoft transcribes speech before the speaker finishes
- Tavus limits Griffin’s live video conversations to a research preview
- Decagon Voice 3 keeps the conversation going while tasks run
- Anthropic puts enterprise AI engineers through a 12-week project residency
- Barclays expands Claude across engineering and banking operations
- Apple will require explicit approval for broad Mac file access
Columns
- BitNest’s speed needs an action-time denominator
- Financing moves first on the rural data-center map
- Giving the rules engine the final count in a compliance agent
- Do AstaBrief’s citations preserve a finding’s scope?
- TopK-Guided measures its gain in operations
- Codex makes the environment part of the task
- For a small business, Muse starts with the connections
- Gimlet’s speed target depends on a data-center launch
- Holo4 puts cross-interface control in the developer’s hands
- Who held decision authority in Atria’s task logs?
- Synthesia’s journalist avatar stays inside one story
- AIHW found no breach, and the traces show attempts
- OpenAI cannot tie 53 images to the people who uploaded them
- GitHub's fuzzing agent writes the harness and leaves each crash for a person to read
- EvasionBench folds the continue prompt into the score
- An open tab is where Gemini in Chrome builds a quiz
- Oracle separates Jupiter's 2028 date from the payment duty
- Gemini 3.8 Flash TTS ties consent to the 30-second sample
- NADI's 40 million euros fund chip-design teams
- The cheaper Opus call leaves the check on the desk
- MiMo-V2.6 opens the weights and moves the bottleneck to hardware
- Muse hit its first real gate in shopping
- 7,000 sales do not measure productive capacity
- Step 5's two parameter counts measure two different loads
- Does Muse have a switch that turns Memory off?
- RoboHarm measures whether a text refusal survives on a robot arm
- When the instruction file changes, which layer still holds the lock?
- Claude Science timed runtime, not binder quality
- Google CC shares the household; a child still has no account
- 30,000 agents and a 6 percent safety share, still no watt