Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
81 changes: 74 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

<div align="center">

<b>LLM inference in C/C++</b>
<b>LLM inference in C/C++, with batched constrained decisions over text and images</b>

[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://opensource.org/licenses/MIT)
[![Release](https://img.shields.io/github/v/release/ggml-org/llama.cpp?filter=v*&color=brightgreen)](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0)
Expand All @@ -17,7 +17,74 @@

</div>

## Quick start
## Constrained decisions

Instead of generating a JSON object one token at a time, this branch scores a whole schema in a single batched
forward pass. Every field has a fixed set of allowed values, so the values are scored as token paths that fork from
the same KV cache. All fields are answered in one `llama_decode`, cannot see each other, and the object is
assembled by code, so the output always matches the schema. Each field comes back with a probability.

Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and chunks encode through
the same mtmd path the completion endpoint uses. A decision over a screenshot, a document scan, or a whole folder of
images stays a single pass rather than becoming a run of generate calls.

```bash
./build/bin/llama-server -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \
--decision-seqs 8 --host 0.0.0.0 --port 8081
```

```bash
curl http://localhost:8081/decision -H "Content-Type: application/json" -d '{
"instructions": "Answer each question about this screenshot.",
"schema": {
"properties": {
"page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]},
"error": {"type": "boolean"}
}
},
"contexts": ["What kind of page is this?"],
"images": ["iVBORw0KGgoAAAANSUhEUg..."]
}'
```

```json
{
"object": "decision",
"results": [
{
"decision": {"page": "settings", "error": true},
"fields": {
"page": {"value": "settings", "probability": 0.868, "scored_nodes": 1, "tree": true},
"error": {"value": true, "probability": 0.966, "scored_nodes": 1, "tree": true}
},
"usage": {"context_tokens": 516, "scored_rows": 9}
}
]
}
```

`images` is positional: entry *i* belongs to context *i*. An entry is one base64 string, or an array when a single
context should see several images. Media markers already present in the context text are left where the caller put
them, so images can be interleaved with the caller's own labels.

Full reference: [parallel-decision](tools/parallel-decision/README.md).

### Vision decision harness

`tools/parallel-decision/examples/vision-decision-harness/` is a runnable example UI for the endpoint. It does
folder upload, batch runs over images and text with SSE progress, image selection across a folder, a text
classification suite, and snippet export. It is an example, not a dependency, and nothing in the server links
against it. See its [README](tools/parallel-decision/examples/vision-decision-harness/README.md).

### Set `LLAMA_DECISION_DEBUG`

Traces tokenization, chunk encoding, and decode on the decision path.

## Standard llama.cpp

Everything below this point is the upstream project as usual.

### Quick start

A few options to get `llama.cpp` installed on your machine:

Expand Down Expand Up @@ -49,7 +116,7 @@ llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
</tr>
<table>

## Description
### Description

The main goal of `llama.cpp` is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
Expand All @@ -65,7 +132,7 @@ a wide range of hardware - locally and in the cloud.

The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-org/ggml) library.

## Supported backends
### Supported backends

| Backend | Target devices |
| --- | --- |
Expand All @@ -87,7 +154,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or
| [WebGPU](docs/build.md#webgpu) | All |
| [ZenDNN](docs/build.md#zendnn) | AMD CPU |

## Documentation
### Documentation

#### Tools

Expand All @@ -109,15 +176,15 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or
- [Models](docs/models.md)
- [Release process](docs/release.md)

## Contributing
### Contributing

- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the `llama.cpp` repo and merge PRs into the `master` branch
- Any help with managing issues, PRs and projects is very appreciated!
- Read the [CONTRIBUTING.md](CONTRIBUTING.md) for more information

## Acknowledgements
### Acknowledgements

- [yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) - Single-header HTTP server, used by `llama-server` - MIT license
- [nothings/stb](https://github.com/nothings/stb) - Single-header image format decoder, used by multimodal subsystem - Public domain
Expand Down
5 changes: 5 additions & 0 deletions tools/mtmd/mtmd.h
Original file line number Diff line number Diff line change
Expand Up @@ -507,6 +507,11 @@ struct bitmap {
struct bitmaps {
std::vector<bitmap> entries;
~bitmaps() = default;
bitmaps() = default;
bitmaps(bitmaps && other) noexcept = default;
bitmaps & operator=(bitmaps && other) noexcept = default;
bitmaps(const bitmaps &) = delete;
bitmaps & operator=(const bitmaps &) = delete;
// return list of pointers to mtmd_bitmap
// example:
// auto bitmaps_c_ptr = bitmaps.c_ptr();
Expand Down
2 changes: 1 addition & 1 deletion tools/parallel-decision/CMakeLists.txt
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# shared engine: used by llama-parallel-decision and by llama-server's /decision endpoint
add_library(llama-decision STATIC decision-engine.cpp decision-engine.h)
target_include_directories(llama-decision PUBLIC ${CMAKE_CURRENT_SOURCE_DIR})
target_link_libraries(llama-decision PUBLIC llama-common llama)
target_link_libraries(llama-decision PUBLIC llama-common llama mtmd)
target_compile_features(llama-decision PUBLIC cxx_std_17)
# llama-server links it into a shared library (libllama-server-impl)
set_target_properties(llama-decision PROPERTIES POSITION_INDEPENDENT_CODE ON)
Expand Down
72 changes: 69 additions & 3 deletions tools/parallel-decision/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,12 @@ After the context, each field's allowed values are scored as token paths that fo
fields are answered in one `llama_decode` and cannot see each other. Each answer comes back with a probability, and
the JSON object is assembled by code, so it always matches the schema.

Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and the chunks are encoded
through the same mtmd path the completion endpoint uses, so a decision runs over a screenshot, a document scan, or a
folder of images without leaving the single-pass scoring model.

This directory holds the engine (`decision-engine.*`), a CLI (`llama-parallel-decision`), and the engine is also
served by `llama-server` as `POST /v1/decision`.
served by `llama-server` as `POST /decision`.

## Build

Expand Down Expand Up @@ -59,13 +63,13 @@ sequence (about 50 MB each for Qwen3.5 4B and 9B), and llama.cpp only batches th
the same number of tokens. The engine right-pads each group of branches to its longest one, so they still score in a
single pass; the padding comes after the token that is read, so it doesn't change the result.

## POST /v1/decision
## POST /decision

`contexts` is a list of 1-256 strings. They share one schema, one set of instructions, and one cached prefix; results
come back in the same order.

```bash
curl http://localhost:8096/v1/decision -H "Content-Type: application/json" -d '{
curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{
"model": "gemma-4-12b",
"instructions": "Answer each question about this support request from its state.",
"schema": {
Expand Down Expand Up @@ -122,6 +126,35 @@ Numeric fields take `aggregate`: `mode` (default), `median` or `mean`.
| `tree_max` | 128 | per-field switch between tree and greedy |
| `cache_prompt` | true | reuse the cached instructions + schema prefix |

## Images

Add `images` alongside `contexts`. It is positional: entry *i* belongs to context *i*. An entry is either one
base64 string or an array of base64 strings when a single context should see several images. A data URL prefix
is accepted and stripped.

```bash
curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{
"instructions": "Answer each question about this screenshot.",
"schema": {
"properties": {
"page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]},
"error": {"type": "boolean"}
}
},
"contexts": ["What kind of page is this?"],
"images": ["iVBORw0KGgoAAAANSUhEUg..."]
}'
```

Images are placed by the media marker that `mtmd` reserves for the loaded projector. If the context text already
contains that marker, the marker is left where the caller put it, so you can interleave markers with your own labels
and control which part of the text each image belongs to. If the text has no marker, one is prepended per image.

Because the chunks are encoded in the same batched pass that scores the branches, adding images does not turn a
decision into a sequence of generate calls.

Set `LLAMA_DECISION_DEBUG` in the environment to trace tokenization, chunk encoding, and decode on this path.

## CLI

`llama-parallel-decision` runs the same engine from a worker process (stdin/stdout protocol, one JSON request per
Expand All @@ -132,3 +165,36 @@ line). Environment: `DECIDE_TREE`, `DECIDE_TREE_MAX`, `DECIDE_NSEQ`, `DECIDE_SPL
[decision-playground](https://github.com/thecodacus/decision-playground) is a browser-only playground: it talks
straight to your llama-server, runs a decision and the same question as a chat completion side by side with live
timers, and has a small game whose agents decide through the endpoint.

### Vision Decision Harness (Example)

For multimodal decision testing with image support, this directory ships a runnable example under
`examples/vision-decision-harness/`. It is a small Flask web UI that:

- Accepts image folder uploads and runs them through `/decision` in a batch
- Adds an image selection mode: every image in a folder is scored in one decision pass and the
UI reports the single best match for a question
- Streams results back over SSE as each file is scored
- Ships a text test suite for measuring classification accuracy and calibration
- Ships `tests/scan_for_secrets.py`, which uses the decision endpoint itself to flag files that
look like they contain private data before you commit

Build and run the server first:

```bash
./build/bin/llama-server --host 0.0.0.0 --port 8081 \
-m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \
--decision-seqs 8
```

Then the harness:

```bash
cd tools/parallel-decision/examples/vision-decision-harness
pip install -r requirements.txt
python3 app.py
```

The harness listens on port 5786 and proxies to the server on port 8081. Set `LLAMA_SERVER_URL`
to point it somewhere else.

Loading