New information at the request boundary

An application using llama.cpp can now learn more than a model’s name from its server. Version 0.6.0 adds an architecture object describing supported input and output modalities to the model listing. That changes the boundary encountered by a client preparing an image-bearing request: the software can query the kinds of input the backend accepts. The same release extends the embedding interface to typed image, audio and video inputs. An embedding is a numerical representation of data, so that request path serves a different purpose from asking for a conversational response.[1]

Capability discovery moves an integration decision from a model-name list maintained inside the application towards a conversation with the server. A client can base its input path on reported capabilities instead of inferring them from a name. The useful change is a more explicit reason for choosing a particular request path. It also requires the client to consume the new fields: an older client that ignores them retains its own decision logic. The practical boundary of model switching therefore lies in adapting to the capability contract as well as selecting a different name.[1]

The data path inside

A second change sits beneath the incoming server request. The extended batch interface carries tokens and embedding vectors together and allows additional state per token. Examples, speculative decoding, multimodal tools and the server have migrated to that interface. The project’s own clients have therefore moved onto the new data path, alongside the changes at external endpoints. A developer reading the release can follow a concrete integration boundary: how accepted content becomes processing input. That is a change in the software plumbing that applications depend on when they submit different kinds of content.[1]

For the team owning an integration, the work extends from choosing a model to owning the route taken by its inputs. Server capability fields and internal interface changes touch different stages of the same request. An application using the ready-made server can concentrate on its external contract; a team integrating the library also has to connect the batch interface to its own code. Keeping control over more layers distributes review work across those layers. This is a tangible division of responsibility between consuming a service and maintaining the local machinery that executes a request.[1]

Where the session resumes

The request path continues after an answer is produced: session state can be saved and restored later. This release raises the session-format version to 11 and the sequence-state version to 4. Fixes cover cleanup after failed restoration and mismatched cache rotation, while the server rejects truncated multimodal input. These changes concern the handling of saved state and incomplete requests as well as model execution. The release identifies specific points where a local application hands information back to the inference system, including the attempt to resume an earlier session.[1]

Control over local inference is consequently tied to a concrete maintenance responsibility. When one team owns the input path, server capabilities and saved state, a version change brings all three into its migration decision. The workload can be narrower for an application using the standard server without persisting sessions. My decision rule from this release is to identify which interfaces and state formats an application owns when choosing local execution, alongside the models it can run. Clear ownership puts the work supplied by the release and the maintenance carried by the team into the same request path.[1]