Eigen RadarAI
Analysis

DeepSeek adds sight while Nvidia moves capability into the harness and cache

DeepSeek gave an experimental Flash model image understanding. Nvidia researchers separately raised ARC-AGI-3 performance with a surrounding harness and mapped one model's prompt cache into another without retraining either model.

Artificial Intelligence··Midday
On a research bench, a camera, processor modules and an irregular memory bridge connecting them.

DeepSeek Flash reads images through the API

DeepSeek released the experimental DeepSeek-V4-Flash-Vision-Exp model on its developer platform, adding image understanding to V4-Flash. Access is through the API and the weights are not open. The company says text performance is unchanged and multimodal-agent scores approach Opus-4.8, though no independent test supports that comparison. Images are billed as tokens, with a maximum of 384 per image, and coverage also reported a limit of 600 images per request.[1]

Nvidia raises the ARC-AGI-3 result with a harness

Nvidia's Agentic Variation Operators research puts memory, tools, and a supervisor around the existing Claude Opus 5 model. On ARC-AGI-3, a benchmark made of two-dimensional games without instructions, the model scored 30 percent alone and 100 percent with the harness. Nvidia argues that open harnesses give teams more settings that can affect accuracy. TechCrunch also notes that OpenAI tripled its own result by changing two harness settings, without reaching Nvidia's level.[2]

One model's prompt memory transfers to another

Nvidia researchers mapped the key-value cache built while one model reads a prompt into a second model with a linear transform. Moving a 32,768-token cache from Qwen3 14B to 32B took 278 ms, compared with 7 seconds for standard re-prefilling. Transfers ran 2.7 to 25 times faster and retained up to 98 percent of the target model's standalone accuracy. The linear method failed on Ministral pairs, where a nonlinear network was needed to recover above 90 percent.[3]

References

  1. News sourceDeepSeekDeepSeek's Flash model learns to read screens↩
  2. News sourceTechCrunchNvidia's scaffolding took a model from 30 to 100 on ARC-AGI-3↩
  3. News sourceVentureBeatA small model's memory can be handed to a big one with plain algebra↩