Eigen RadarAI
Analysis

Cohere makes cost per page its edge in document parsing

Cohere made Parse 5 generally available for turning PDFs and images into structured text, pricing 1,000 pages at 1.5 dollars while its own benchmark places higher-scoring rivals ahead. NVIDIA's TensorRT Model Connect takes open models into a C++ inference bundle in two commands. The developments clarify the practical choice between buying document processing as a service and placing a model in a team's own runtime.

Artificial Intelligence··Night
Two hands feed a crumpled blank page into an unbranded scanner on a wooden workbench as clean blank cards emerge in orderly layers into a shallow tray.

Parse 5 puts cost first

Cohere made Parse 5, a 2.3 billion parameter model that turns PDFs, slides and images into structured Markdown, generally available. It prices 1,000 pages at 1.5 dollars. On Cohere's own ParseBench ranking, Parse 5 scores 79.2, behind GPT-5.5 at 84.4 and Opus 4.8 at 84.3. Those scores have not been independently tested. The model is served through the Cohere API and Model Vault as well as Microsoft Foundry and AWS SageMaker.[1]

NVIDIA moves a model into C++ in two steps

NVIDIA's TensorRT Model Connect brings together reference implementations for open models. A first step on the Python command line turns a Hugging Face model identifier or local checkpoint into a deployment bundle. A C++ runtime then loads that bundle without requiring a PyTorch or Python installation. The tool supports more than 80 model families, including Nemotron Speech and Qwen 3 VL, and its source code is available in NVIDIA's public repository.[2]

The deployment path follows the use case

Parse 5 is priced by the page as a ready service that preserves document structure, while TensorRT Model Connect carries a model bundle into a team's own C++ application. Cohere offers a clear unit cost for teams that want document parsing handled as one service. NVIDIA separates a module interface that reaches tensors and components directly from a semantic interface that accepts prompts, images and audio. Custom GPU kernels can be added without rebuilding the application, and the project publishes nightly releases.[1], [2]

References

  1. News sourceVentureBeatCohere's new document model trails on the benchmark and leads on cost per page↩1↩2
  2. News sourceNVIDIA Technical BlogNvidia publishes a tool that takes an open model to C++ inference in two commands↩1↩2