Eigen RadarAI
Analysis

TensorRT Edge-LLM finishes Thor's agent benchmark 6.4 times faster

NVIDIA says TensorRT Edge-LLM ran Qwen3.6-27B on one Jetson AGX Thor for the MLPerf Inference v6.1 Edge Agentic workload. AI Daily Post repeats 24 minutes 36 seconds versus 2 hours 37 minutes for llama.cpp, 6.4 times, at 52.33 tokens per second across 1,007 turns. Those timings sit in both accounts.

Artificial Intelligence··Evening
An unmarked compact edge-compute kit with a spinning cooling fan sits on a sunlit electronics workbench, surrounded by cables and unmarked components; a different idle box is blurred behind it.

24 minutes 36 seconds on one Thor

NVIDIA’s developer blog says TensorRT Edge-LLM completed the MLPerf Inference v6.1 Edge Agentic benchmark 6.4 times faster on Jetson AGX Thor. AI Daily Post, edited by Brian Petersen and dated 16 September, puts the TensorRT run at 24 minutes 36 seconds on a single Jetson AGX Thor Developer Kit, against 2 hours 37 minutes for the llama.cpp reference on the same hardware and the same Qwen3.6-27B model. It calls that a 6.4 times gap. The 24-minute and 2-hour-37-minute prints are in the AI Daily Post copy; NVIDIA’s pool title names the 6.4 times Thor result without repeating those clock times.[1], [2]

1,007 turns at 52.33 tokens per second

AI Daily Post says the system ran at 52.33 tokens per second across all 1,007 turns in the workload. It says Edge Agentic replays software-engineering agent trajectories: 20 conversations, 1,007 turns, with input length climbing to roughly 23.5K tokens, rather than scoring a single prompt-response pair. NVIDIA’s 16 September blog is the first-party MLPerf write-up those figures accompany. The 20-conversation shape is the workload AI Daily Post printed.[1], [2]

Same kit, two stacks

AI Daily Post underlines that the 6.4 times comparison is on the same hardware and the same Qwen3.6-27B model. NVIDIA presents the TensorRT Edge-LLM Thor submission as the faster stack. Neither text, in the material used here, converts that laboratory SingleStream result into a robot or vehicle deployment claim.[1], [2]

References

  1. News sourceNVIDIANVIDIA published TensorRT Edge-LLM's MLPerf Edge Agentic time on Jetson AGX Thor↩1↩2↩3
  2. News sourceAI Daily PostTensorRT Edge-LLM finishes 1,007 agent turns in 24 minutes 36 seconds↩1↩2↩3