Azure and AWS are centralising agent controls at gateways, while Anthropic's self-hosted option and a framework comparison expose the operational choices behind enterprise deployments.
Artificial Intelligence··Midday
The gateway moves beyond individual calls
Microsoft has put a separate AI gateway tier for Azure API Management into public preview, organising models and tools apart from conventional APIs. In the system described by InfoQ, OpenAI, Anthropic and Mistral models can sit on the same control plane as connections to AWS Bedrock and Google Vertex AI. Remote MCP servers, OpenAPI specifications and built-in connectors reaching more than 1,000 SaaS applications can also be gathered behind the gateway. Administrators can set token and request limits, quotas, content safety and model fallback in a portal; the resource provisions in about a minute without scale-unit planning. The tier is not yet generally available. AWS, meanwhile, has extended control in the Bedrock AgentCore gateway from individual calls to sequences of actions within a session. Temporal policies can compare a value with an earlier call's result, tally session spending, require a prescribed order or demand human approval before a significant action. Per-user rate limits cover requests, tokens and connection time. AWS says enforcement sits outside agent code, and rate limiting needs no code change.[1], [3]
The boundary of self-hosted execution
Anthropic's self-hosted environments, now in public beta, allow Claude Code sessions to run on machines supplied by the customer. The option is available for Team and Enterprise plans and is disabled by default. According to the announcement, repository checkouts, build artifacts and files created during a session remain in the organisation's infrastructure. Prompts, responses, tool results and session transcripts still go to Anthropic, however, and code that Claude reads leaves the environment as prompt content. Self-hosted execution therefore does not close every data path inside the customer boundary. The fact that organisations using the zero-data-retention setting cannot enable the feature makes that limit especially relevant. The operating burden also belongs to the customer: Anthropic says companies need engineers to prepare runner images, update the infrastructure and run an orchestration layer for scaling on demand. Location consequently becomes a question of who provisions, maintains and expands the environment as well as where files remain. An organisation gains more direct control over the execution setting, while still having to assess the service's data connection and the additional work required to keep sessions available.[2]
The runtime layer changes the result
The importance of gateway and execution choices becomes tangible in a Composio comparison reported by The Decoder. The company ran DeepSeek V4 Flash through Claude Code, Codex, OpenCode and Oh My Pi on 30 tasks involving real tools such as Gmail, GitHub, Slack and Notion. Success rates were close, but time and cost diverged. Oh My Pi completed 17 of 30 tasks and was slowest at 272 seconds per task. OpenCode succeeded on 14 of 30 and was cheapest at 0.073 dollars per successful task. Claude Code was fastest at 122 seconds and most expensive at 0.195 dollars per successful task. On 7 tasks, success or failure changed solely with the framework. Composio produced the measurement and published it through X, and there is no independent replication. The scope is limited to the selected 30 tasks and the published measurement. Even at that scale, the comparison helps explain why enterprises must treat the gateway, agent framework, data path and operating model as separate choices alongside the model itself. Azure and AWS controls determine which calls proceed, in what order and within which limits. Anthropic's option separates where execution occurs from which data reaches the service. The Composio results show that the same model can still produce different task outcomes, costs and completion times across runtime layers.[4], [1], [3], [2]