llama.cpp Introduces Decision Models
huggingface llama
| Source: Mastodon | Original article
llama.cpp now adds support for decision models, enabling local AI inference through its server’s /v1/system endpoint.
Llama.cpp, the open‑source C/C++ engine that has become the go‑to tool for running large language models locally, has added a new “decision model” capability. The update introduces a /v1/systemone endpoint that accepts a state – which can be plain text, JSON data or even a screenshot – together with a set of typed questions. In a single forward pass the model returns a probability for each option, rather than generating token‑by‑token text as traditional chat models do.
The feature mirrors the System One format first described by TypeSafe’s Jev model, and it brings that workflow onto users’ own hardware. By handling the entire decision in one inference step, developers can achieve lower latency and deterministic outputs, making llama.cpp a viable alternative for software automation, classification, and other decision‑oriented AI tasks that previously required cloud‑hosted services.
Why it matters for the Nordic AI ecosystem is twofold. First, the ability to run decision models locally aligns with the region’s strong emphasis on data privacy and edge computing, allowing enterprises to keep sensitive inputs on‑premise without sacrificing model performance. Second, the lightweight nature of llama.cpp – optimized for ARM, Apple Silicon, AVX‑512 and other architectures – means the new API can be deployed on a wide range of devices, from laptops to embedded systems, potentially lowering the cost barrier for AI‑driven decision making.
What to watch next are the early adopters’ performance reports and any expansion of the model catalog beyond the initial Jev‑style offerings. The community will also be looking for tooling that integrates the /v1/systemone endpoint into existing pipelines, as well as benchmarks that compare local decision inference against cloud alternatives. As the feature matures, it could reshape how Nordic developers embed AI reasoning directly into their products.
Sources
Back to AIPULSEN