GameWAM predicts frame and keystroke together as Apple builds tests from MCP
A preprint published on arXiv on 29 August introduced GameWAM, a world action model that generates the next visual frame and keyboard-and-mouse trajectory together. Apple researchers described Agent Seer, which builds agent test scenarios from Model Context Protocol specifications; parameter schema complexity drives quality most. One line offers parallel action prediction at the interface; the other automated tests from tool definitions.
Artificial Intelligence··Midday
GameWAM produces frame and input in parallel
A preprint posted to arXiv on 29 August presents GameWAM, which the authors describe as the first world action model built for playing games and driving a graphical interface through the same loop. It generates the next visual frame and the matching keyboard-and-mouse trajectory in parallel, then replans from fresh observations in blocks so that long sessions stay on course. The authors report task success comparable to the baselines they measure against while executing fewer native actions, and they describe a failure mode they call Low-Frequency Action Source Imprinting, in which the model becomes sensitive to where an action came from. The 44-page paper is a preprint on arXiv and has not been peer-reviewed, and the authors position parallel frame-and-input prediction as the mechanism that keeps long game or interface sessions aligned without restarting the loop from scratch after every keystroke. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1]
Agent Seer generates scenarios from MCP definitions
Apple researchers describe Agent Seer, which builds test scenarios for tool-using agents from Model Context Protocol specifications instead of hand-written cases. The system enriches function names, descriptions and parameter schemas, produces graded scenarios with synthetic outputs and turns them into multi-turn dialogues. Tested on 7 separate MCP specifications, the pipeline covered every tool on small and medium specifications. The team reports that parameter schema complexity drives quality more than tool-suite size, which comes second. In imperfect scenarios the main failure mode is picking the wrong argument values, a distinction coarse-grained metrics usually miss. The pipeline is presented as a way to generate graded scenarios with synthetic outputs rather than relying on a fixed library of hand-written prompts for each new tool definition a product team publishes. The approach is aimed at teams that publish MCP tool definitions faster than manual test libraries can be maintained. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[2]
Prediction and tests meet two separate agent needs
GameWAM models the next frame and input together on a graphical interface, enabling replanning across long sessions; the authors report similar success with fewer native actions, yet Action Source Imprinting shows sensitivity to where an action came from. Agent Seer generates tests from tool definitions to check whether functions connected through MCP are called with correct arguments, and it covered every tool on small and medium MCP specifications in the team's reported runs. The two works target different agent layers: one perception and action generation in games and graphical interfaces, the other tool-call accuracy when schemas grow complex. Apple's team says wrong argument values dominate imperfect scenarios, while GameWAM's authors warn that low-frequency action sources can imprint on the model. For readers the concrete development is both an interface-level world model and protocol-level tests announced in the same week. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1], [2]