Skip to content
 
 

Latest commit

 

History

11,056 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama.cpp

llama

LLM inference in C/C++, with batched constrained decisions over text and images

License: MIT Release Nightly Server Docker Winget

ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API

Constrained decisions

Instead of generating a JSON object one token at a time, this branch scores a whole schema in a single batched forward pass. Every field has a fixed set of allowed values, so the values are scored as token paths that fork from the same KV cache. All fields are answered in one llama_decode, cannot see each other, and the object is assembled by code, so the output always matches the schema. Each field comes back with a probability.

Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and chunks encode through the same mtmd path the completion endpoint uses. A decision over a screenshot, a document scan, or a whole folder of images stays a single pass rather than becoming a run of generate calls.

./build/bin/llama-server -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \
    --decision-seqs 8 --host 0.0.0.0 --port 8081
curl http://localhost:8081/decision -H "Content-Type: application/json" -d '{
  "instructions": "Answer each question about this screenshot.",
  "schema": {
    "properties": {
      "page":  {"type": "string", "enum": ["login", "checkout", "settings", "other"]},
      "error": {"type": "boolean"}
    }
  },
  "contexts": ["What kind of page is this?"],
  "images": ["iVBORw0KGgoAAAANSUhEUg..."]
}'
{
  "object": "decision",
  "results": [
    {
      "decision": {"page": "settings", "error": true},
      "fields": {
        "page":  {"value": "settings", "probability": 0.868, "scored_nodes": 1, "tree": true},
        "error": {"value": true,      "probability": 0.966, "scored_nodes": 1, "tree": true}
      },
      "usage": {"context_tokens": 516, "scored_rows": 9}
    }
  ]
}

images is positional: entry i belongs to context i. An entry is one base64 string, or an array when a single context should see several images. Media markers already present in the context text are left where the caller put them, so images can be interleaved with the caller's own labels.

Full reference: parallel-decision.

Vision decision harness

tools/parallel-decision/examples/vision-decision-harness/ is a runnable example UI for the endpoint. It does folder upload, batch runs over images and text with SSE progress, image selection across a folder, a text classification suite, and snippet export. It is an example, not a dependency, and nothing in the server links against it. See its README.

Set LLAMA_DECISION_DEBUG

Traces tokenization, chunk encoding, and decode on the decision path.

Standard llama.cpp

Everything below this point is the upstream project as usual.

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

LLM inference in C/C++

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages