Strata adds parallel conversations and a local connection for Codex
Strata v0.1.39 lets its local language-model engine handle simultaneous conversations and exposes a Responses API connection for Codex CLI. Requests wait when the configured conversation limit is reached, while each active conversation reduces memory available to the model’s expert cache. The release’s tested tool loop has specific unsupported features. Experimental hardware paths also carry stated testing limits, so the new connections and build options do not establish identical support on every computer.
Artificial Intelligence··Night
Local conversations can run together
A Strata session no longer has to wait for the previous conversation to finish before another can run. LinuxCompatible reports that version 0.1.39 adds parallel conversations, replacing version 0.1.38’s one-request-at-a-time handling. New requests wait for a free slot once the configured limit is reached. There is a memory cost: each conversation takes space otherwise available to the model’s expert cache. Parallel operation is therefore recommended when most of those expert components fit in graphics memory.[1], [2]
Codex gets a local connection
The release also offers a Responses API endpoint, a software connection that lets Codex CLI use a local instance. Testing with version 0.160.0 of Codex included a loop that invokes tools. Support has boundaries: identifiers for earlier responses, hosted tools and reasoning summaries are not available. A user connecting a local model therefore receives a defined subset of the interface, rather than every feature associated with the remote service.[1]
Hardware options carry testing limits
Some additional hardware paths remain experimental. LinuxCompatible describes older and experimental options that were compiled and checked without testing on the corresponding physical devices. The update also includes an optional multi-GPU expert-offload change that affects numerical rounding. These stated limits accompany the release’s broader local-use options.[1]
Long inputs also have a changed processing path. The release notes say that the engine now chooses a chunk size automatically and budgets its streamed expert ring in bytes. A lengthy prompt can give different outputs from version 0.1.38 when expert components pass through different cache groups; an option restores the earlier ring behavior. Hardware additions include a separate build using CUDA 12.9 for Pascal and Volta NVIDIA cards. The Intel Arc path requires a source build on Linux. The developer says that this version was compiled and its kernel tests run on a CPU device, while the corresponding Arc hardware was not tested. Source builds for older processors without AVX2 are also offered with physical-device testing limits. These are the developer’s stated capabilities and checks, not an independent reproduction of all hardware results.[2], [1]