Cohere makes cost per page its edge in document parsing
Cohere made Parse 5 generally available for turning PDFs and images into structured text, pricing 1,000 pages at 1.5 dollars while its own benchmark places higher-scoring rivals ahead. NVIDIA's TensorRT Model Connect takes open models into a C++ inference bundle in two commands. The developments clarify the practical choice between buying document processing as a service and placing a model in a team's own runtime.
Artificial Intelligence··Night
Parse 5 puts cost first
Cohere made Parse 5, a 2.3 billion parameter model that turns PDFs, slides and images into structured Markdown, generally available. It prices 1,000 pages at 1.5 dollars. On Cohere's own ParseBench ranking, Parse 5 scores 79.2, behind GPT-5.5 at 84.4 and Opus 4.8 at 84.3. Those scores have not been independently tested. The model is served through the Cohere API and Model Vault as well as Microsoft Foundry and AWS SageMaker.[1]
NVIDIA moves a model into C++ in two steps
NVIDIA's TensorRT Model Connect brings together reference implementations for open models. A first step on the Python command line turns a Hugging Face model identifier or local checkpoint into a deployment bundle. A C++ runtime then loads that bundle without requiring a PyTorch or Python installation. The tool supports more than 80 model families, including Nemotron Speech and Qwen 3 VL, and its source code is available in NVIDIA's public repository.[2]
The deployment path follows the use case
Parse 5 is priced by the page as a ready service that preserves document structure, while TensorRT Model Connect carries a model bundle into a team's own C++ application. Cohere offers a clear unit cost for teams that want document parsing handled as one service. NVIDIA separates a module interface that reaches tensors and components directly from a semantic interface that accepts prompts, images and audio. Custom GPU kernels can be added without rebuilding the application, and the project publishes nightly releases.[1], [2]