Eigen RadarAI
Analysis

The agent production stack is splitting into three layers

Presence packages enterprise deployment with engineers, Embabel exposes orchestration in code, and Meta adds a separate memory agent; together they show why long-task reliability cannot be reduced to one model feature.

Artificial Intelligence··Morning
Synthetic system linking a long agent workflow emerging from a violet deployment corridor to an overhead memory loop

Presence seeks reliability in a service package

OpenAI introduced Presence as an enterprise offering for moving agents into production in customer service and internal workflows. It builds on the company's customisable Workspace Agents and is currently available only to qualifying enterprise customers. Workflow selection, system integration, guideline definition and production testing are not left to software documentation alone: OpenAI's forward deployed engineers provide tailored support through those stages. The disclosed arrangement therefore treats reliability as a property of setup and operations as well as model output. No price was published, however, and the report gives no measurements that would let an outside team reproduce the same arrangement in its own environment. THE DECODER also says it remains unclear how specific compliance requirements such as the European Union's AI Act will be handled. The visible boundary of the Presence announcement is consequently precise: access and engineering support are described, while comparable outcomes such as setup time, the share of human review or successful rollback are not. The offer places part of production reliability in a service layer, but the available information does not yet make that layer independently measurable outside a customer deployment.[1]

Embabel exposes orchestration as code

Embabel 1.0 approaches the same production problem through a downloadable Java framework. Co-created by Spring Framework founder Rod Johnson, it represents an agent as typed domain objects made of goals, actions and connecting conditions rather than as a manually ordered sequence of prompts and tool calls. Built on Spring AI, it can work with providers including OpenAI, Anthropic, Gemini, Bedrock, Mistral and DeepSeek, as well as local endpoints such as Ollama. Goal-oriented planning determines the action order at runtime, and the plan can be reassessed when conditions change during a task. Individual actions can be routed to a particular model, model types can be represented by role aliases, and planning can be combined with explicit state machines inside the same agent. The visible layer here is orchestration logic that can be read and run against different task sets, rather than a customer-specific service process. That openness does not by itself establish better performance; the report contains no same-task comparison between Embabel and Presence. It does make the objects available for examination different. Presence offers a supported deployment, whereas Embabel presents the relationships among goals, actions and conditions as software structures a developer can inspect and execute.[2]

Meta separates task history into another agent

A setup published by a Meta AI team addresses a third layer: preserving state across long tasks. Under behavioural state decay, the action agent forgets constraints, repeats failed commands and rediscovers earlier errors. The second agent does not perform the task. It updates a structured memory bank of status, knowledge and procedure, then decides whether to add a reminder to the action agent's next call. First-attempt success on Terminal-Bench 2.0 rises from 38 percent to 46 percent, while the task-weighted Tau2-Bench average rises from 55 percent to 62 percent. The gains vary by domain: roughly 10 points for airline and retail tasks and 3 points for telecom. Claude Opus 4.6 served as the memory agent, while Claude Sonnet 4.5 and Qwen3.5-27B were action agents. The arXiv paper, accompanied by code, has not been peer reviewed. Read together, the three reports place reliability at distinct intervention points: enterprise support, executable orchestration and task memory. They do not compare the approaches in one experiment and establish no ranking. Their common finding is narrower: long-task reliability depends on which layer accompanies the model and what measurement is disclosed for it.[3], [1], [2]

References

  1. News sourceTHE DECODEROpenAI introduced Presence for enterprise agent deployment↩1↩2
  2. News sourceInfoQThe Java agent framework Embabel reached 1.0↩1↩2
  3. News sourceTHE DECODERMeta researchers are testing a second, memory-keeping agent for long tasks↩