DeepSeek adds sight while Nvidia moves capability into the harness and cache
DeepSeek gave an experimental Flash model image understanding. Nvidia researchers separately raised ARC-AGI-3 performance with a surrounding harness and mapped one model's prompt cache into another without retraining either model.
Artificial Intelligence··Midday
DeepSeek Flash reads images through the API
DeepSeek released the experimental DeepSeek-V4-Flash-Vision-Exp model on its developer platform, adding image understanding to V4-Flash. Access is through the API and the weights are not open. The company says text performance is unchanged and multimodal-agent scores approach Opus-4.8, though no independent test supports that comparison. Images are billed as tokens, with a maximum of 384 per image, and coverage also reported a limit of 600 images per request.[1]
Nvidia raises the ARC-AGI-3 result with a harness
Nvidia's Agentic Variation Operators research puts memory, tools, and a supervisor around the existing Claude Opus 5 model. On ARC-AGI-3, a benchmark made of two-dimensional games without instructions, the model scored 30 percent alone and 100 percent with the harness. Nvidia argues that open harnesses give teams more settings that can affect accuracy. TechCrunch also notes that OpenAI tripled its own result by changing two harness settings, without reaching Nvidia's level.[2]
One model's prompt memory transfers to another
Nvidia researchers mapped the key-value cache built while one model reads a prompt into a second model with a linear transform. Moving a 32,768-token cache from Qwen3 14B to 32B took 278 ms, compared with 7 seconds for standard re-prefilling. Transfers ran 2.7 to 25 times faster and retained up to 98 percent of the target model's standalone accuracy. The linear method failed on Ministral pairs, where a nonlinear network was needed to recover above 90 percent.[3]