Visual Studio Copilot gains a thinking-effort dial as Apple auto-builds agent tests
With GitHub's August update, organisations can publish custom Copilot agents across repositories and have Visual Studio detect them automatically, and developers can set the model's thinking effort to low, medium or high. Apple researchers described Agent Seer, which builds agent test scenarios from Model Context Protocol specifications and reports that parameter schema complexity drives quality more than tool-suite size. Both companies are easing how software agents are built and tested.
Artificial Intelligence··Morning
Organisation-wide custom agents open in Visual Studio
With GitHub's August update, organisations can publish custom Copilot agents across repositories and have Visual Studio detect them automatically. A developer can set the model's thinking effort to low, medium or high, trading reasoning depth against token consumption. The update also brings usage details and plan information into the context window and warns as a limit approaches. Favourite models can be pinned and unused ones collapsed, with model capabilities, context window sizes and cost information in the same place.[1]
The Git agent reviews work before a pull request
The Git agent reviews uncommitted changes and commits before a pull request is opened and shows its findings inline in the editor. The features are available on the Free, Student, Pro, Pro+, Max, Business and Enterprise plans. The update builds on the earlier step of bringing organisation-wide custom agents into repositories; developers can now both share an agent and receive automated review before sending code.[1]
Apple generates tests from MCP specifications
Apple researchers describe Agent Seer, which builds test scenarios for tool-using agents from Model Context Protocol specifications instead of hand-written cases. The system enriches function names, descriptions and parameter schemas, produces graded scenarios with synthetic outputs and turns them into multi-turn dialogues. Tested on 7 separate MCP specifications, the pipeline covered every tool on small and medium specifications. The team reports that parameter schema complexity drives quality more than tool-suite size, which comes second. In imperfect scenarios the main failure mode is picking the wrong argument values.[2]