Setapp access and llamafile packages for local inference
MacPaw and Liquid AI plan to open local inference to Setapp developers, while Mozilla.ai packages transcription binaries and two local model builds as llamafile v0.10.5 release artifacts.
Artificial Intelligence··Evening
The planned local inference route through Setapp
MacPaw is working with Liquid AI on locally hosted models for its own products and plans to open that capability to developers publishing through Setapp, its app store with more than 150000 paying users. MacPaw is also building a locally hosted version of its Eney assistant. Liquid AI provides Elix, an on-device inference system paired with a local memory component, and tailors models to the hardware. MacPaw chief executive Oleksandr Kosovan says local models will allow assistants and agentic workflows to run offline. Liquid AI chief executive Ramin Hasani describes the work as a customization stack around models, where models can use data and improve. Setapp is testing credit-based pricing that deducts credits according to task complexity. The report gives no release date, developer pricing or list of supported devices. The announced route therefore defines a plan for developer access through the store and the local inference components behind it, while timing, device coverage and commercial terms have not yet been published.[1]
Ready-made artifacts in the llamafile release
Mozilla.ai's llamafile v0.10.5 release presents a different distribution format for local use. It follows three llama.cpp synchronizations completed within two weeks and adds prebuilt transcribefile speech-to-text binaries to the release artifacts. Expanded documentation covers the help system, differences among release binaries and graphics processor support, including the Vulkan backend. The release notes also highlight two model builds. Ternary Bonsai 27B is a compressed Qwen3.6-27B model using ternary weights at 1.58 bits per weight; it occupies about 6 GB on disk and works multimodally with an optional vision tower. Laguna-S-2.1 is poolside's 118B mixture-of-experts coding model; it runs 8B active parameters per token and is configured in this package for a 256K context window. Mozilla.ai says both builds run at usable speeds on consumer hardware with enough memory. It publishes no throughput or latency measurement supporting that statement. The notes also credit community contributions to the transcribefile artifacts and documentation work.[2]
Different scopes and delivery stages
The reports differ in scope and delivery stage. The MacPaw-Liquid AI route is a planned Setapp developer capability with no announced date or developer price; Mozilla.ai's route is an already released llamafile package. In the first case, local inference is intended to expand from Eney, Elix and a local memory component prepared for MacPaw products toward Setapp developers. In the llamafile release, distribution takes the form of ready-made transcription binaries, graphics processor documentation and two named model builds. The products cover different model families and tasks. Their common ground is the creation of packaging and access paths that let developers run inference on local hardware. Setapp's schedule, price and device coverage remain unpublished; the Mozilla.ai usable-speed statement lacks throughput or latency measurements. One route represents planned store access and the other current release artifacts, giving developers concrete and distinct channels for local inference.[1], [2]