Eigen RadarAI
Analysis

Cloudflare traces agents, researchers reject their papers, and AWS counts their tokens

Cloudflare has added agent tracing, an AWS example bolts on token telemetry, and researchers rejected two agent-written papers. Faster autonomous workflows still leave operators with a heavy verification and observability burden.

Artificial Intelligence··Evening
In a bright operations room, three luminous data paths feed a transparent tracing rig while a reviewer examines them and two blank pages return.

Agent traces and opposite privacy defaults at Cloudflare

According to InfoQ, Cloudflare has opened the first part of Cloudflare Agents, a dashboard that collects traces of deployed agent sessions with spans for model calls, tool runs and approvals. What is stored depends on the library a team uses: Think and wrapAISDK() keep no message or tool payloads unless a flag is set, while Flue stores messages, system instructions, tool definitions, arguments and results by default. Traces carry agent name, agent ID and conversation ID, and session replay shows messages, reasoning, tool calls and subagent activity across turns, though not images. The feature is free in beta and moves to Workers Observability pricing on 1 October 2026, with 200,000 events a day and 3 days of retention on the free tier and 20 million events a month with 7 days of retention on the paid tier, then 0.60 dollars per further million. InfoQ notes that traces are not lossless: long messages can hit span size limits, and approval spans record lifecycle events rather than how long a person waited. Observability is productised, yet operators still must check which library stores what and which gaps remain.[1]

Two agent-written papers rejected on NeurIPS questions

In the Shadow Evaluation study reported by The Decoder, researchers from Princeton and the UK AI Security Institute gave AI agents unpublished NeurIPS 2026 research questions, then had the original authors review the results as peer reviewers. Both agent-written papers were rejected, one with a strong reject. Claude Opus 4.8 with extra-high reasoning and GPT-5.6 Sol each had six days, 3,000 dollars in API credits, a GPU budget and virtual-machine access, and spent about 1,130 dollars of the credit. The agents handled the engineering work but abandoned ambitious goals within 10 hours of a 36 to 48 hour allocation, exceeded paper length limits and drifted from their instructions. Reviewers criticised poorly motivated data, unreadable prose and bizarre experiment choices. The sample is two NeurIPS submissions, so the result speaks to those tasks rather than research in general; the write-up is at cruxevals.com. Set against claims that autonomous research is near, the trial puts engineering capacity and peer-review standards on the same table and shows that speed does not guarantee acceptance.[2]

Token telemetry bolted on in an AWS example

An AWS Machine Learning Blog post describes a multi-agent system built with Strands Agents, in which an orchestrator routes requests to a budget agent and a financial analysis agent and the whole system is deployed to the Bedrock AgentCore runtime. Models come from Qwen 3.5 9B on SageMaker AI through a vLLM container, and Claude Haiku 4.5 and Claude Sonnet 4.6 through Bedrock. SageMaker's OpenAI-compatible endpoints produce no automatic token telemetry, so the authors emit OpenTelemetry gen_ai.chat spans by hand and switch on include_usage inside stream_options so that vLLM reports token counts. Authentication runs through a custom httpx.Auth subclass with auto-refreshing bearer tokens. Traces land in CloudWatch Transaction Search and AWS X-Ray; the model runs on ml.g6e.2xlarge instances with a vLLM container at version 0.22.1, and the code sits in aws-samples/sagemaker-genai-hosting-examples. Read with Cloudflare's dashboard and the Shadow Evaluation rejections, the picture is plain: as agent workflows speed up, tracing, cost and scientific verification stay on the operator. Productised monitoring still depends on library defaults; the sample build bolts telemetry on later; peer-review standards remain under human oversight.[3], [1], [2]

References

  1. News sourceInfoQAgent tracing arrives at Cloudflare with opposite privacy defaults↩1↩2
  2. News sourceThe DecoderTwo agent-written papers were rejected by the researchers who set the questions↩1↩2
  3. News sourceAWS Machine Learning BlogAWS walks through a multi-agent build that has to add its own token telemetry↩