diff --git a/README.md b/README.md index aae3bcd35ad9..6570d55cc56f 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,7 @@
-LLM inference in C/C++ +LLM inference in C/C++, with batched constrained decisions over text and images [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](https://opensource.org/licenses/MIT) [![Release](https://img.shields.io/github/v/release/ggml-org/llama.cpp?filter=v*&color=brightgreen)](https://github.com/ggml-org/llama.cpp/releases?q=tag:v0) @@ -17,7 +17,74 @@
-## Quick start +## Constrained decisions + +Instead of generating a JSON object one token at a time, this branch scores a whole schema in a single batched +forward pass. Every field has a fixed set of allowed values, so the values are scored as token paths that fork from +the same KV cache. All fields are answered in one `llama_decode`, cannot see each other, and the object is +assembled by code, so the output always matches the schema. Each field comes back with a probability. + +Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and chunks encode through +the same mtmd path the completion endpoint uses. A decision over a screenshot, a document scan, or a whole folder of +images stays a single pass rather than becoming a run of generate calls. + +```bash +./build/bin/llama-server -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \ + --decision-seqs 8 --host 0.0.0.0 --port 8081 +``` + +```bash +curl http://localhost:8081/decision -H "Content-Type: application/json" -d '{ + "instructions": "Answer each question about this screenshot.", + "schema": { + "properties": { + "page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]}, + "error": {"type": "boolean"} + } + }, + "contexts": ["What kind of page is this?"], + "images": ["iVBORw0KGgoAAAANSUhEUg..."] +}' +``` + +```json +{ + "object": "decision", + "results": [ + { + "decision": {"page": "settings", "error": true}, + "fields": { + "page": {"value": "settings", "probability": 0.868, "scored_nodes": 1, "tree": true}, + "error": {"value": true, "probability": 0.966, "scored_nodes": 1, "tree": true} + }, + "usage": {"context_tokens": 516, "scored_rows": 9} + } + ] +} +``` + +`images` is positional: entry *i* belongs to context *i*. An entry is one base64 string, or an array when a single +context should see several images. Media markers already present in the context text are left where the caller put +them, so images can be interleaved with the caller's own labels. + +Full reference: [parallel-decision](tools/parallel-decision/README.md). + +### Vision decision harness + +`tools/parallel-decision/examples/vision-decision-harness/` is a runnable example UI for the endpoint. It does +folder upload, batch runs over images and text with SSE progress, image selection across a folder, a text +classification suite, and snippet export. It is an example, not a dependency, and nothing in the server links +against it. See its [README](tools/parallel-decision/examples/vision-decision-harness/README.md). + +### Set `LLAMA_DECISION_DEBUG` + +Traces tokenization, chunk encoding, and decode on the decision path. + +## Standard llama.cpp + +Everything below this point is the upstream project as usual. + +### Quick start A few options to get `llama.cpp` installed on your machine: @@ -49,7 +116,7 @@ llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF -## Description +### Description The main goal of `llama.cpp` is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. @@ -65,7 +132,7 @@ a wide range of hardware - locally and in the cloud. The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-org/ggml) library. -## Supported backends +### Supported backends | Backend | Target devices | | --- | --- | @@ -87,7 +154,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or | [WebGPU](docs/build.md#webgpu) | All | | [ZenDNN](docs/build.md#zendnn) | AMD CPU | -## Documentation +### Documentation #### Tools @@ -109,7 +176,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or - [Models](docs/models.md) - [Release process](docs/release.md) -## Contributing +### Contributing - Contributors can open PRs - Collaborators will be invited based on contributions @@ -117,7 +184,7 @@ The `llama.cpp` project is build on top of the [ggml](https://github.com/ggml-or - Any help with managing issues, PRs and projects is very appreciated! - Read the [CONTRIBUTING.md](CONTRIBUTING.md) for more information -## Acknowledgements +### Acknowledgements - [yhirose/cpp-httplib](https://github.com/yhirose/cpp-httplib) - Single-header HTTP server, used by `llama-server` - MIT license - [nothings/stb](https://github.com/nothings/stb) - Single-header image format decoder, used by multimodal subsystem - Public domain diff --git a/tools/mtmd/mtmd.h b/tools/mtmd/mtmd.h index c2de26eeee2c..85ca905afaff 100644 --- a/tools/mtmd/mtmd.h +++ b/tools/mtmd/mtmd.h @@ -507,6 +507,11 @@ struct bitmap { struct bitmaps { std::vector entries; ~bitmaps() = default; + bitmaps() = default; + bitmaps(bitmaps && other) noexcept = default; + bitmaps & operator=(bitmaps && other) noexcept = default; + bitmaps(const bitmaps &) = delete; + bitmaps & operator=(const bitmaps &) = delete; // return list of pointers to mtmd_bitmap // example: // auto bitmaps_c_ptr = bitmaps.c_ptr(); diff --git a/tools/parallel-decision/CMakeLists.txt b/tools/parallel-decision/CMakeLists.txt index f4d41d2f27e9..2c7ca8d68e02 100644 --- a/tools/parallel-decision/CMakeLists.txt +++ b/tools/parallel-decision/CMakeLists.txt @@ -1,7 +1,7 @@ # shared engine: used by llama-parallel-decision and by llama-server's /decision endpoint add_library(llama-decision STATIC decision-engine.cpp decision-engine.h) target_include_directories(llama-decision PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}) -target_link_libraries(llama-decision PUBLIC llama-common llama) +target_link_libraries(llama-decision PUBLIC llama-common llama mtmd) target_compile_features(llama-decision PUBLIC cxx_std_17) # llama-server links it into a shared library (libllama-server-impl) set_target_properties(llama-decision PROPERTIES POSITION_INDEPENDENT_CODE ON) diff --git a/tools/parallel-decision/README.md b/tools/parallel-decision/README.md index 76e86b90e773..27e37f0ffad0 100644 --- a/tools/parallel-decision/README.md +++ b/tools/parallel-decision/README.md @@ -7,8 +7,12 @@ After the context, each field's allowed values are scored as token paths that fo fields are answered in one `llama_decode` and cannot see each other. Each answer comes back with a probability, and the JSON object is assembled by code, so it always matches the schema. +Contexts can carry images. A prompt part is either a run of text tokens or a media chunk, and the chunks are encoded +through the same mtmd path the completion endpoint uses, so a decision runs over a screenshot, a document scan, or a +folder of images without leaving the single-pass scoring model. + This directory holds the engine (`decision-engine.*`), a CLI (`llama-parallel-decision`), and the engine is also -served by `llama-server` as `POST /v1/decision`. +served by `llama-server` as `POST /decision`. ## Build @@ -59,13 +63,13 @@ sequence (about 50 MB each for Qwen3.5 4B and 9B), and llama.cpp only batches th the same number of tokens. The engine right-pads each group of branches to its longest one, so they still score in a single pass; the padding comes after the token that is read, so it doesn't change the result. -## POST /v1/decision +## POST /decision `contexts` is a list of 1-256 strings. They share one schema, one set of instructions, and one cached prefix; results come back in the same order. ```bash -curl http://localhost:8096/v1/decision -H "Content-Type: application/json" -d '{ +curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{ "model": "gemma-4-12b", "instructions": "Answer each question about this support request from its state.", "schema": { @@ -122,6 +126,35 @@ Numeric fields take `aggregate`: `mode` (default), `median` or `mean`. | `tree_max` | 128 | per-field switch between tree and greedy | | `cache_prompt` | true | reuse the cached instructions + schema prefix | +## Images + +Add `images` alongside `contexts`. It is positional: entry *i* belongs to context *i*. An entry is either one +base64 string or an array of base64 strings when a single context should see several images. A data URL prefix +is accepted and stripped. + +```bash +curl http://localhost:8096/decision -H "Content-Type: application/json" -d '{ + "instructions": "Answer each question about this screenshot.", + "schema": { + "properties": { + "page": {"type": "string", "enum": ["login", "checkout", "settings", "other"]}, + "error": {"type": "boolean"} + } + }, + "contexts": ["What kind of page is this?"], + "images": ["iVBORw0KGgoAAAANSUhEUg..."] +}' +``` + +Images are placed by the media marker that `mtmd` reserves for the loaded projector. If the context text already +contains that marker, the marker is left where the caller put it, so you can interleave markers with your own labels +and control which part of the text each image belongs to. If the text has no marker, one is prepended per image. + +Because the chunks are encoded in the same batched pass that scores the branches, adding images does not turn a +decision into a sequence of generate calls. + +Set `LLAMA_DECISION_DEBUG` in the environment to trace tokenization, chunk encoding, and decode on this path. + ## CLI `llama-parallel-decision` runs the same engine from a worker process (stdin/stdout protocol, one JSON request per @@ -132,3 +165,36 @@ line). Environment: `DECIDE_TREE`, `DECIDE_TREE_MAX`, `DECIDE_NSEQ`, `DECIDE_SPL [decision-playground](https://github.com/thecodacus/decision-playground) is a browser-only playground: it talks straight to your llama-server, runs a decision and the same question as a chat completion side by side with live timers, and has a small game whose agents decide through the endpoint. + +### Vision Decision Harness (Example) + +For multimodal decision testing with image support, this directory ships a runnable example under +`examples/vision-decision-harness/`. It is a small Flask web UI that: + +- Accepts image folder uploads and runs them through `/decision` in a batch +- Adds an image selection mode: every image in a folder is scored in one decision pass and the + UI reports the single best match for a question +- Streams results back over SSE as each file is scored +- Ships a text test suite for measuring classification accuracy and calibration +- Ships `tests/scan_for_secrets.py`, which uses the decision endpoint itself to flag files that + look like they contain private data before you commit + +Build and run the server first: + +```bash +./build/bin/llama-server --host 0.0.0.0 --port 8081 \ + -m model-Q4_K_M.gguf --mmproj mmproj-model.gguf \ + --decision-seqs 8 +``` + +Then the harness: + +```bash +cd tools/parallel-decision/examples/vision-decision-harness +pip install -r requirements.txt +python3 app.py +``` + +The harness listens on port 5786 and proxies to the server on port 8081. Set `LLAMA_SERVER_URL` +to point it somewhere else. + diff --git a/tools/parallel-decision/decision-engine.cpp b/tools/parallel-decision/decision-engine.cpp index c0f1d65000c5..cf630b942368 100644 --- a/tools/parallel-decision/decision-engine.cpp +++ b/tools/parallel-decision/decision-engine.cpp @@ -2,6 +2,7 @@ #include "chat.h" #include "common.h" +#include "mtmd-helper.h" #include #include @@ -10,6 +11,10 @@ #include #include +// Toggle debug output with LLAMA_DECISION_DEBUG env var +namespace { bool debug_enabled() { static bool v = std::getenv("LLAMA_DECISION_DEBUG") != nullptr; return v; } } +#define DECISION_DEBUG(fmt, ...) do { if (debug_enabled()) { fprintf(stderr, "[decision-debug] " fmt "\n", ##__VA_ARGS__); } } while(0) + namespace llama_decision { namespace { @@ -173,10 +178,11 @@ struct decision_field { // ---------------------------------------------------------------- engine -engine::engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs) +engine::engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs, mtmd_context * mctx) : ctx(ctx), vocab(llama_model_get_vocab(llama_get_model(ctx))), mem(llama_get_memory(ctx)), seq_snap(seq_base), seq_pool(seq_base + 1), n_pool(n_seqs - 1), - pad_branches(llama_model_is_recurrent(llama_get_model(ctx)) || llama_model_is_hybrid(llama_get_model(ctx))) { + pad_branches(llama_model_is_recurrent(llama_get_model(ctx)) || llama_model_is_hybrid(llama_get_model(ctx))), + mctx(mctx) { if (n_seqs < 3) { throw std::invalid_argument("a decision engine needs at least 3 sequences"); } @@ -192,11 +198,94 @@ tokens_t engine::tokenize(const std::string & text, bool add_special) const { return toks; } +// Tokenize text that may contain media markers. When mctx is available and the text +// contains a media marker, uses mtmd_tokenize to split the text into text chunks and +// image chunks. Text tokens with LLAMA_TOKEN_NULL at image positions are returned, +// and the image chunks are collected for separate encoding in decode_parts. +engine::multimodal_tokens engine::tokenize_mm(const std::string & text, bool add_special, + const mtmd::bitmaps * bitmaps) const { + multimodal_tokens result; + + if (!mctx) { + // no multimodal context, fall back to regular tokenization + result.toks = tokenize(text, add_special); + return result; + } + + const char * marker = mctx ? mtmd_get_marker(mctx) : nullptr; + if (marker == nullptr || text.find(marker) == std::string::npos) { + // no media marker in text, use regular tokenization + result.toks = tokenize(text, add_special); + return result; + } + + // Use mtmd_tokenize to properly split text and image chunks, passing bitmaps + mtmd::input_chunks chunks(mtmd_input_chunks_init()); // MUST be initialized! + mtmd_input_text input_text; + input_text.text = text.c_str(); + input_text.text_len = text.size(); + input_text.add_special = add_special; + input_text.parse_special = true; + + auto bmp_ptr = [bitmaps]() { + std::vector res; + if (bitmaps) { + res.reserve(bitmaps->entries.size()); + for (const auto & b : bitmaps->entries) { + res.push_back(b.ptr.get()); + } + } + return res; + }(); + int32_t rc = mtmd_tokenize(mctx, chunks.ptr.get(), &input_text, bmp_ptr.data(), (int32_t) bmp_ptr.size()); + DECISION_DEBUG("tokenize_mm: mtmd_tokenize rc=%d chunks.size=%zu", (int)rc, chunks.size()); + if (rc != 0) { + // fall back to regular tokenization on error + result.toks = tokenize(text, add_special); + return result; + } + + // Extract tokens from chunks, interleaving text tokens and LLAMA_TOKEN_NULL for images + for (size_t i = 0; i < chunks.size(); ++i) { + const mtmd_input_chunk * chunk = chunks[i]; + enum mtmd_input_chunk_type type = mtmd_input_chunk_get_type(chunk); + if (type == MTMD_INPUT_CHUNK_TYPE_TEXT) { + size_t n_tokens = 0; + const llama_token * toks = mtmd_input_chunk_get_tokens_text(chunk, &n_tokens); + for (size_t j = 0; j < n_tokens; ++j) { + result.toks.push_back(toks[j]); + } + } else if (type == MTMD_INPUT_CHUNK_TYPE_IMAGE) { + // image chunk: insert LLAMA_TOKEN_NULL placeholder and record the chunk copy + size_t n_tokens = mtmd_input_chunk_get_n_tokens(chunk); + result.toks.insert(result.toks.end(), n_tokens, LLAMA_TOKEN_NULL); + // copy the chunk so it stays valid after chunks object is destroyed + result.chunks.push_back(mtmd_input_chunk_copy(chunk)); + } else if (type == MTMD_INPUT_CHUNK_TYPE_AUDIO) { + // audio chunk: insert LLAMA_TOKEN_NULL placeholder and record the chunk copy + size_t n_tokens = mtmd_input_chunk_get_n_tokens(chunk); + result.toks.insert(result.toks.end(), n_tokens, LLAMA_TOKEN_NULL); + result.chunks.push_back(mtmd_input_chunk_copy(chunk)); + } + } + + // deduplicate BOS if needed (consistent with tokenize()) + const llama_token bos = llama_vocab_bos(vocab); + if (result.toks.size() >= 2 && result.toks[0] == bos && result.toks[1] == bos) { + result.toks.erase(result.toks.begin()); + } + + return result; +} + // Decode several prompts, each on its own sequence, packed into as few batches as n_batch allows. +// Image chunks are encoded via the mtmd batch API (mtmd_batch_init/add_chunk/encode/get_output_embd) +// followed by mtmd_helper_decode_image_chunk, the same path the server's completion endpoint uses. +// Text tokens use a plain llama_batch for batched llama_decode. void engine::decode_parts(const std::vector & parts) { const int n_batch = (int) llama_n_batch(ctx); llama_batch batch = llama_batch_init(n_batch, 0, 1); - auto flush = [&]() { + auto flush_text = [&]() { const int rc = batch.n_tokens > 0 ? llama_decode(ctx, batch) : 0; common_batch_clear(batch); if (rc != 0) { @@ -205,15 +294,93 @@ void engine::decode_parts(const std::vector & parts) { : "llama_decode failed on the decision prompt (" + std::to_string(rc) + ")"); } }; + // Collect image chunks that need encoding + std::vector img_chunks; + for (const auto & p : parts) { + if (p.kind == prompt_part::CHUNK) { + img_chunks.push_back(p.chunk); + } + } + + // Encode image chunks via the mtmd batch API, the same path the server's completion + // endpoint uses. Some models (e.g. SmolVLM2) do not support batching in the CLIP context, + // so mtmd_batch_add_chunk rejects the second chunk with "batch too large" (rc=2). In that + // case each chunk is encoded on its own with mtmd_encode_chunk instead. + // Text-only decisions never reach this and skip the whole block. + mtmd::batch_ptr mbatch; + bool use_batch = false; + if (!img_chunks.empty() && mctx) { + mbatch.reset(mtmd_batch_init(mctx)); + use_batch = true; + for (auto * chunk : img_chunks) { + DECISION_DEBUG("decode_parts: adding image chunk with n_tokens=%d", (int) mtmd_input_chunk_get_n_tokens(chunk)); + int32_t add_rc = mtmd_batch_add_chunk(mbatch.get(), chunk); + DECISION_DEBUG("decode_parts: mtmd_batch_add_chunk rc=%d", (int) add_rc); + if (add_rc == 2) { + // batch too large: this model does not support batched clip encoding + use_batch = false; + mbatch.reset(); + break; + } + if (add_rc != 0) { + throw std::runtime_error("mtmd_batch_add_chunk failed (" + std::to_string(add_rc) + ")"); + } + } + if (use_batch) { + int32_t enc_rc = mtmd_batch_encode(mbatch.get()); + DECISION_DEBUG("decode_parts: mtmd_batch_encode rc=%d", (int) enc_rc); + if (enc_rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_batch_encode failed on the decision prompt (" + std::to_string(enc_rc) + ")"); + } + } + } + + // Now iterate parts: decode text via llama_batch, decode image via mtmd_helper_decode_image_chunk for (const auto & p : parts) { - for (size_t i = 0; i < p.toks->size(); ++i) { - if (batch.n_tokens == n_batch) { - flush(); + if (p.kind == prompt_part::CHUNK) { + // Flush pending text tokens first so image decoding starts at the right position + flush_text(); + DECISION_DEBUG("decode_parts: decoding CHUNK pos0=%d seq=%d", (int) p.pos0, (int) p.seq); + float * embd = nullptr; + if (use_batch && mbatch) { + embd = mtmd_batch_get_output_embd(mbatch.get(), p.chunk); + DECISION_DEBUG("decode_parts: embd ptr from batch=%p", (void*)embd); + } else { + // Non-batch fallback: encode this chunk individually + DECISION_DEBUG("decode_parts: encoding chunk individually via mtmd_encode_chunk"); + int32_t enc_rc = mtmd_encode_chunk(mctx, p.chunk); + DECISION_DEBUG("decode_parts: mtmd_encode_chunk rc=%d", (int) enc_rc); + if (enc_rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_encode_chunk failed on the decision prompt (" + std::to_string(enc_rc) + ")"); + } + embd = mtmd_get_output_embd(mctx); + DECISION_DEBUG("decode_parts: embd ptr from mtmd=%p", (void*)embd); + } + if (!embd) { + llama_batch_free(batch); + throw std::runtime_error("failed to get image embedding for chunk"); + } + DECISION_DEBUG("decode_parts: mtmd_helper_decode_image_chunk n_batch=%d", n_batch); + llama_pos new_n_past = p.pos0; + int32_t rc = mtmd_helper_decode_image_chunk(mctx, ctx, p.chunk, embd, p.pos0, p.seq, n_batch, &new_n_past, nullptr, nullptr); + DECISION_DEBUG("decode_parts: mtmd_helper_decode_image_chunk rc=%d new_n_past=%d", (int) rc, (int) new_n_past); + if (rc != 0) { + llama_batch_free(batch); + throw std::runtime_error("mtmd_helper_decode_image_chunk failed on the decision prompt (" + std::to_string(rc) + ")"); + } + } else { + DECISION_DEBUG("decode_parts: TOKS pos0=%d seq=%d n_toks=%zu", (int) p.pos0, (int) p.seq, p.toks->size()); + for (size_t i = 0; i < p.toks->size(); ++i) { + if (batch.n_tokens == n_batch) { + flush_text(); + } + common_batch_add(batch, (*p.toks)[i], p.pos0 + (llama_pos) i, { p.seq }, false); } - common_batch_add(batch, (*p.toks)[i], p.pos0 + (llama_pos) i, { p.seq }, false); } } - flush(); + flush_text(); llama_batch_free(batch); } @@ -229,7 +396,9 @@ bool engine::prepare_prefix(const tokens_t & shared, bool allow_cache) { } cached.clear(); if (!shared.empty()) { - decode_parts({ { &shared, 0, seq_snap } }); + DECISION_DEBUG("prepare_prefix: calling decode_parts with shared.size=%zu", shared.size()); + decode_parts({ { prompt_part::TOKS, &shared, nullptr, 0, seq_snap } }); + DECISION_DEBUG("prepare_prefix: decode_parts returned"); cached = shared; } return false; @@ -326,12 +495,15 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve throw std::invalid_argument("a decision needs at least one context"); } const tokens_t shared = tokenize(shared_text, true); - std::vector prefixes; - for (const auto & text : contexts) { - prefixes.push_back(tokenize(text, shared.empty())); - if (prefixes.back().empty()) { + std::vector prefixes; + for (size_t i = 0; i < contexts.size(); ++i) { + const mtmd::bitmaps * ctx_bitmaps = (i < opt.context_bitmaps.size()) ? &opt.context_bitmaps[i] : nullptr; + auto mtoks = tokenize_mm(contexts[i], shared.empty(), ctx_bitmaps); + DECISION_DEBUG("decide_batch: tokenize_mm returned toks=%zu chunks=%zu", mtoks.toks.size(), mtoks.chunks.size()); + if (mtoks.toks.empty()) { throw std::invalid_argument("the decision context must not be empty"); } + prefixes.push_back(std::move(mtoks)); } std::vector fields; @@ -406,6 +578,16 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve const size_t n_group = std::min(per_group, contexts.size() - g0); const auto tp = std::chrono::steady_clock::now(); + // We need to keep segment token vectors alive while parts reference them + // Reserve enough capacity to prevent reallocation (which would invalidate pointers) + // With N image chunks interleaved with text, there can be up to N+1 text segments. + // Use a generous reserve based on the max chunks in any prefix. + std::vector seg_storage; + size_t max_chunks = 0; + for (const auto & mtoks : prefixes) { + max_chunks = std::max(max_chunks, mtoks.chunks.size()); + } + seg_storage.reserve((max_chunks + 1) * n_group); // at most (chunks+1) text segments per context std::vector parts; for (size_t i = 0; i < n_group; ++i) { const llama_seq_id trunk = seq_pool + (llama_seq_id) i; @@ -413,9 +595,44 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve if (!shared.empty()) { llama_memory_seq_cp(mem, seq_snap, trunk, -1, -1); } - parts.push_back({ &prefixes[g0 + i], (llama_pos) shared.size(), trunk }); + // Build prompt_parts from multimodal_tokens: text tokens become TOKS parts, + // image chunks become CHUNK parts with proper position tracking + const auto & mtoks = prefixes[g0 + i]; + llama_pos pos = (llama_pos) shared.size(); + size_t chunk_idx = 0; + size_t seg_start = 0; + for (size_t t_idx = 0; t_idx < mtoks.toks.size(); ++t_idx) { + if (mtoks.toks[t_idx] == LLAMA_TOKEN_NULL) { + // Flush preceding text tokens as a TOKS part + if (t_idx > seg_start) { + seg_storage.push_back(tokens_t(mtoks.toks.begin() + seg_start, mtoks.toks.begin() + t_idx)); + parts.push_back({ prompt_part::TOKS, &seg_storage.back(), nullptr, pos, trunk }); + pos += (llama_pos) seg_storage.back().size(); + seg_start = t_idx + 1; + } + // Add image chunk if available + if (chunk_idx < mtoks.chunks.size()) { + const size_t n_img_tokens = mtmd_input_chunk_get_n_tokens(mtoks.chunks[chunk_idx]); + parts.push_back({ prompt_part::CHUNK, nullptr, mtoks.chunks[chunk_idx], pos, trunk }); + pos += (llama_pos) n_img_tokens; + chunk_idx++; + seg_start = t_idx + 1; + } else { + // No image chunk is left for this NULL token, so skip it + // (it's a placeholder that has no corresponding chunk) + seg_start = t_idx + 1; + } + } + } + // Flush remaining text tokens + if (seg_start < mtoks.toks.size()) { + seg_storage.push_back(tokens_t(mtoks.toks.begin() + seg_start, mtoks.toks.end())); + parts.push_back({ prompt_part::TOKS, &seg_storage.back(), nullptr, pos, trunk }); + } } + DECISION_DEBUG("decide_batch: calling decode_parts with %zu parts", parts.size()); decode_parts(parts); + DECISION_DEBUG("decide_batch: decode_parts returned"); llama_synchronize(ctx); // llama_decode is asynchronous: wait for the prefill so its time isn't billed to scoring out.prefill_ms += ms_since(tp); @@ -429,7 +646,7 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve std::vector> owner; // (context in group, field) for (size_t i = 0; i < n_group; ++i) { const llama_seq_id trunk = seq_pool + (llama_seq_id) i; - const llama_pos pos0 = (llama_pos) (shared.size() + prefixes[g0 + i].size()); + const llama_pos pos0 = (llama_pos) (shared.size() + prefixes[g0 + i].toks.size()); for (size_t f = 0; f < state[i].size(); ++f) { auto & fd = state[i][f]; if (fd.use_tree) { @@ -487,7 +704,7 @@ batch_result engine::decide_batch(const std::string & shared_text, const std::ve for (size_t i = 0; i < n_group; ++i) { llama_memory_seq_rm(mem, seq_pool + (llama_seq_id) i, -1, -1); result & r = out.items[g0 + i]; - r.context_tokens = prefixes[g0 + i].size(); + r.context_tokens = prefixes[g0 + i].toks.size(); r.rows = total; for (auto & fd : state[i]) { if (fd.use_tree && fd.probs.empty()) { diff --git a/tools/parallel-decision/decision-engine.h b/tools/parallel-decision/decision-engine.h index 49e7f104fb44..40ddfc6dbf54 100644 --- a/tools/parallel-decision/decision-engine.h +++ b/tools/parallel-decision/decision-engine.h @@ -11,6 +11,8 @@ // larger fields walk the trie greedily. #include "llama.h" +#include "mtmd.h" +#include "mtmd-helper.h" #include "json.h" #include @@ -34,6 +36,10 @@ struct options { size_t tree_max = 128; bool split_boundary = false; // legacy: tokenise suffix and values separately bool allow_cache = true; // reuse the cached static prefix when it matches + // Optional per-context bitmaps for multimodal decision. If non-empty, must have the same + // size as the contexts vector in decide_batch. Each entry holds bitmaps to prepend to + // that context (the context text should contain media markers at the corresponding positions). + std::vector context_bitmaps; }; struct field_result { @@ -72,7 +78,7 @@ struct batch_result { // flight, then branches. The context needs a unified KV cache so branches share the trunk's cells. class engine { public: - engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs); + engine(llama_context * ctx, llama_seq_id seq_base, int n_seqs, mtmd_context * mctx = nullptr); result decide(const std::string & shared_text, const std::string & context_text, const std::vector & fields, const options & opt); @@ -83,8 +89,11 @@ class engine { const std::vector & fields, const options & opt); private: + // A prompt part can be either a list of text tokens or a media chunk (image/audio). struct prompt_part { - const tokens_t * toks; + enum type { TOKS, CHUNK } kind; + const tokens_t * toks = nullptr; // when kind == TOKS + const mtmd_input_chunk * chunk = nullptr; // when kind == CHUNK llama_pos pos0; llama_seq_id seq; }; @@ -102,7 +111,40 @@ class engine { int n_pool; bool pad_branches; // recurrent/hybrid model: branches in a decode need equal lengths tokens_t cached; + mtmd_context * mctx; // optional multimodal context for vision input + + // Tokenize text that may contain media markers, expanding them into chunks via mtmd. + // Returns text tokens with LLAMA_TOKEN_NULL at image positions, and the image chunks + // that need to be encoded separately in decode_parts. + struct multimodal_tokens { + tokens_t toks; // text tokens, LLAMA_TOKEN_NULL at image positions + std::vector chunks; // image/audio chunks to decode at those positions + + ~multimodal_tokens() { + for (const auto * chunk : chunks) { + mtmd_input_chunk_free(const_cast(chunk)); + } + } + multimodal_tokens() = default; + multimodal_tokens(multimodal_tokens && other) noexcept + : toks(std::move(other.toks)), chunks(std::move(other.chunks)) {} + multimodal_tokens & operator=(multimodal_tokens && other) noexcept { + if (this != &other) { + for (const auto * chunk : chunks) { + mtmd_input_chunk_free(const_cast(chunk)); + } + toks = std::move(other.toks); + chunks = std::move(other.chunks); + } + return *this; + } + // non-copyable (chunks are owned) + multimodal_tokens(const multimodal_tokens &) = delete; + multimodal_tokens & operator=(const multimodal_tokens &) = delete; + }; + multimodal_tokens tokenize_mm(const std::string & text, bool add_special, + const mtmd::bitmaps * bitmaps = nullptr) const; tokens_t tokenize(const std::string & text, bool add_special) const; void decode_parts(const std::vector & parts); bool prepare_prefix(const tokens_t & shared, bool allow_cache); diff --git a/tools/parallel-decision/examples/vision-decision-harness/.gitignore b/tools/parallel-decision/examples/vision-decision-harness/.gitignore new file mode 100644 index 000000000000..ebbebb5db2a6 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/.gitignore @@ -0,0 +1,7 @@ +__pycache__/ +*.pyc +# written by the test runners on each run +test_results.json +evaluation_results.json +original_jev_script.txt +tests/scan_for_secrets.py diff --git a/tools/parallel-decision/examples/vision-decision-harness/README.md b/tools/parallel-decision/examples/vision-decision-harness/README.md new file mode 100644 index 000000000000..261e2fa57618 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/README.md @@ -0,0 +1,152 @@ +# Vision Decision Harness + +A web-based UI for testing the multimodal `/decision` endpoint of llama-server, built on top of the [parallel-decision branch](https://github.com/thecodacus/llama.cpp) of llama.cpp. + +## Overview + +The harness provides a browser-based interface for evaluating multimodal decision-making capabilities of llama-server. It supports: + +- **Image batch processing**: Upload folders of images and run classification questions +- **Image Selection mode**: Scan a folder of images and find which one matches a question (e.g. "which image contains a rubber duck?") +- **Text file batch processing**: Run text files through the decision endpoint for scoring/classification evaluation +- **Dynamic decision tags**: Add/remove decision options with a + button (tag-list style) +- **Prompt autocompletion**: Get suggestions from the llama-server `/completion` endpoint +- **Benchmarking**: Per-file timing, total batch timing, progress bar +- **Streaming results**: Results appear as each file completes (SSE) +- **Image preview toggle**: Show/hide thumbnails in results +- **Export snippet**: Generate a Python API request snippet with docstring +- **Secret scanning**: Automated scan for sensitive information before committing + +## Prerequisites + +1. **llama-server with decision + multimodal support** — built from the `multimodal-decision` branch of the [thecodacus/llama.cpp](https://github.com/thecodacus/llama.cpp) fork. The server must be started with `--decision-seqs N` (N >= 3) and `--mmproj` for vision support. + + ```bash + ./build/bin/llama-server \ + --host 0.0.0.0 \ + --model model-Q4_K_M.gguf \ + --mmproj mmproj-model.gguf \ + --decision-seqs 8 \ + --device Vulkan1 \ + --port 8081 + ``` + +2. **Python 3** with Flask and requests: + ```bash + pip install flask flask-cors requests + ``` + +## Running + +```bash +cd vision-decision-harness +python3 app.py +``` + +The UI is available at `http://0.0.0.0:5786`. + +The server URL defaults to `http://0.0.0.0:8081`. Override with: +```bash +LLAMA_SERVER_URL=http://localhost:8081 python3 app.py +``` + +## API Endpoints + +| Endpoint | Method | Description | +|---|---|---| +| `/` | GET | Main UI | +| `/api/health` | GET | Health check (proxies to server) | +| `/api/decision` | POST | Proxy to llama-server `/decision` | +| `/api/completion` | POST | Proxy to llama-server `/completion` | +| `/api/upload` | POST | Upload files/folders from browser | +| `/api/batch` | POST | Batch process files (non-streaming) | +| `/api/batch-stream` | POST | Batch process files with SSE streaming | +| `/api/image-selection/stream` | POST | Image selection mode with SSE streaming | +| `/api/export-snippet` | POST | Generate Python API request snippet | +| `/api/tag-suggestions` | POST | Get tag suggestions from `/completion` | + +## Image Selection Mode + +The image selection mode sends all images in a folder to the `/decision` endpoint in a single batched request with a yes/no schema. Each image is evaluated as a separate context, and results are aggregated to identify the single best match. + +**Request:** +```json +{ + "source": "/path/to/image/folder", + "question": "Which image contains a yellow rubber duck?", + "instructions": "Look at this image and determine if it contains the described object. Answer YES if it does, NO if it does not.", + "seed": 42 +} +``` + +**SSE Events:** +- `start`: Found N images, question +- `progress`: Per-image file loaded +- `result`: Per-image yes/no evaluation +- `decision`: Single winning image with file name and probability +- `done`: Total time + +## Text Test Suite + +The `tests/` directory contains a comprehensive text-based decision scoring test suite: + +- `text_samples/` — 20 text files with sentiment classification questions +- `answer_key.json` — Expected results for all test files +- `schema.json` — Shared JSON Schema for the tests +- `evaluate_text_tests.py` — Runner that sends each file through `/decision` and compares to the answer key +- `test_decision_scoring.py` — 41 built-in decision scoring tests (text-only) + +Run with: +```bash +python3 tests/evaluate_text_tests.py +python3 tests/test_decision_scoring.py +``` + +## Secret Scanning + +Before committing, scan all harness files and modified llama.cpp files for sensitive information: + +```bash +python3 tests/scan_for_secrets.py +``` + +This script uses the llama-server `/decision` endpoint to classify each file as containing or not containing: +- Passwords +- API keys +- Email addresses +- Phone numbers +- Home directory paths +- IP addresses +- Credentials/tokens + +Any files flagged as sensitive are reported for manual review before committing. + +## Multimodal Decision Support + +### Server-side Changes + +The fork adds the following to `tools/server/server-context.cpp`: + +1. **Multi-image contexts**: The `images` field can be an array of arrays — one sub-array per context, enabling multiple images per decision context. + +2. **Media marker handling**: Context text can contain inline media markers (fetched from `/props`) to position images at specific points in the text. If no markers are present, they are prepended for backward compatibility. + +3. **Batch vision encoding path**: Uses `mtmd_batch_init` → `mtmd_batch_add_chunk` → `mtmd_batch_encode` → `mtmd_batch_get_output_embd` → `mtmd_helper_decode_image_chunk` for vision encoding, matching the server's working completion path. Falls back to per-chene encoding for models that don't support batch encoding (e.g., SmolVLM2). + +### Decision Engine Changes + +- `decision-engine.h/cpp` — Added `multimodal_tokens` struct, `tokenize_mm()` method, `options::context_bitmaps` field, and batch vision encoding in `decode_parts()` +- `tools/mtmd/mtmd.h` — Added explicit move constructors for `bitmaps` and `bitmap` to support vector storage +- `tools/parallel-decision/CMakeLists.txt` — Links `mtmd` target + +### Debug Toggle + +All debug prints in the decision engine are controlled by the `LLAMA_DECISION_DEBUG` environment variable: + +```bash +LLAMA_DECISION_DEBUG=1 ./build/bin/llama-server --model ... +``` + +## License + +This harness is provided as an example for the parallel-decision multimodal extensions. See the main [llama.cpp README](https://github.com/ggml-org/llama.cpp) for the underlying project license. diff --git a/tools/parallel-decision/examples/vision-decision-harness/app.py b/tools/parallel-decision/examples/vision-decision-harness/app.py new file mode 100644 index 000000000000..3c435eab1803 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/app.py @@ -0,0 +1,507 @@ +#!/usr/bin/env python3 +""" +Vision Decision Harness + +A web UI for testing the multimodal /decision endpoint of llama-server. + +Features: + - Select an image or a folder containing images for batch processing + - Also supports text files in batch (processes .txt and .md files) + - Add decision options dynamically with a + button (tag-list style) + - Prompt autocompletion: type a prompt, and suggestions are fetched + from the llama-server /completion endpoint as you type + - When adding a tag, the harness sends the current prompt to /completion + and uses the model's output to suggest tag names + - Benchmarking: per-file timing, total batch timing, progress bar + - Image Selection mode: scan a folder of images and find which ONE + matches a question (e.g. "which image contains a rubber duck?"). + All images are processed in a SINGLE batched /decision request as + parallel contexts with a yes/no schema. The winning image is + selected from the yes responses, giving ONE result with ONE + probability. + +Runs on 0.0.0.0:5786. Proxies to llama-server at 0.0.0.0:8081 +(override with LLAMA_SERVER_URL env var). +""" + +from flask import Flask, render_template, request, jsonify, Response +from flask_cors import CORS +import requests +import os +import base64 +import json +import time +import uuid +import string + +app = Flask(__name__, static_folder="static", template_folder="templates") +CORS(app) + +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") + +# Temporary upload directory for browser-side file/folder uploads +UPLOAD_DIR = os.environ.get("UPLOAD_DIR", "/tmp/vision-harness-uploads") +os.makedirs(UPLOAD_DIR, exist_ok=True) + + +@app.route("/") +def index(): + return render_template("index.html") + + +@app.route("/api/health") +def health(): + try: + resp = requests.get(f"{SERVER_URL}/v1/models", timeout=5) + return jsonify({"status": "ok", "server": resp.status_code}), 200 + except Exception as e: + return jsonify({"status": "error", "server": str(e)}), 502 + + +@app.route("/api/decision", methods=["POST"]) +def decision(): + data = request.get_json(force=True, silent=True) or {} + try: + resp = requests.post(f"{SERVER_URL}/decision", json=data, timeout=120) + return jsonify(resp.json()), resp.status_code + except Exception as e: + return jsonify({"error": str(e)}), 502 + + +@app.route("/api/completion", methods=["POST"]) +def completion(): + data = request.get_json(force=True, silent=True) or {} + payload = { + "prompt": data.get("prompt", ""), + "n_predict": data.get("n_predict", 16), + "temperature": data.get("temperature", 0.2), + "top_p": data.get("top_p", 0.9), + "seed": data.get("seed", 42), + "stream": False, + } + try: + resp = requests.post(f"{SERVER_URL}/completion", json=payload, timeout=60) + return jsonify(resp.json()), resp.status_code + except Exception as e: + return jsonify({"error": str(e)}), 502 + + +@app.route("/api/upload", methods=["POST"]) +def upload_files(): + """Upload files from the browser (supports folder uploads via webkitdirectory).""" + if "files" not in request.files: + return jsonify({"error": "No files provided"}), 400 + + upload_id = str(uuid.uuid4()) + upload_dir = os.path.join(UPLOAD_DIR, upload_id) + os.makedirs(upload_dir, exist_ok=True) + + uploaded = [] + for storage in request.files.getlist("files"): + rel_path = storage.filename + if not rel_path: + continue + filename = os.path.basename(rel_path) + + # Create subdirectories if this is a folder upload + rel_dir = os.path.dirname(rel_path) + dest_dir = os.path.join(upload_dir, rel_dir) + os.makedirs(dest_dir, exist_ok=True) + dest = os.path.join(dest_dir, filename) + storage.save(dest) + uploaded.append({"path": dest, "name": filename}) + + return jsonify({"upload_id": upload_id, "dir": upload_dir, "files": uploaded, "count": len(uploaded)}) + + +@app.route("/api/batch", methods=["POST"]) +def batch(): + """Batch process all supported files in a folder (or a single file).""" + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + schema = data.get("schema", {}) + instructions = data.get("instructions", "") + contexts = data.get("contexts", [""]) + seed = data.get("seed", 42) + + # Gather files + files = [] + if os.path.isfile(source): + files = [(source, os.path.basename(source))] + elif os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif", ".txt", ".md")): + files.append((fpath, fname)) + + if not files: + return jsonify({"error": "No supported files found", "results": [], "total_files": 0, "total_ms": 0}) + + results = [] + t_start = time.time() + for idx, (fpath, fname) in enumerate(files): + entry = {"file": fname, "index": idx, "total": len(files)} + is_image = fname.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")) + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": seed, + } + if is_image: + try: + with open(fpath, "rb") as f: + img_b64 = base64.b64encode(f.read()).decode() + body["images"] = [img_b64] + except Exception as e: + entry["error"] = f"Failed to read image: {e}" + results.append(entry) + continue + else: + # Text file: use file content as the context + try: + with open(fpath, "r") as f: + body["contexts"] = [f.read()] + except Exception as e: + entry["error"] = f"Failed to read text file: {e}" + results.append(entry) + continue + + t0 = time.time() + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + elapsed = time.time() - t0 + entry["elapsed_ms"] = round(elapsed * 1000, 1) + if resp.status_code == 200: + entry["result"] = resp.json() + else: + entry["error"] = f"HTTP {resp.status_code}: {resp.text[:200]}" + except Exception as e: + entry["error"] = str(e) + results.append(entry) + + total_ms = round((time.time() - t_start) * 1000, 1) + return jsonify({"results": results, "total_files": len(files), "total_ms": total_ms}) + + +@app.route("/api/batch-stream", methods=["POST"]) +def batch_stream(): + """Batch process files and stream results as Server-Sent Events.""" + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + schema = data.get("schema", {}) + instructions = data.get("instructions", "") + contexts = data.get("contexts", [""]) + seed = data.get("seed", 42) + + # Gather files (same logic as /api/batch) + files = [] + if os.path.isfile(source): + files = [(source, os.path.basename(source))] + elif os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif", ".txt", ".md")): + files.append((fpath, fname)) + + def event_stream(): + if not files: + yield "data: " + json.dumps({"error": "No supported files found", "total_files": 0}) + "\n\n" + return + + t_start = time.time() + yield "data: " + json.dumps({"event": "start", "total_files": len(files), "files": [f[1] for f in files]}) + "\n\n" + + for idx, (fpath, fname) in enumerate(files): + t0 = time.time() + entry = {"file": fname, "index": idx, "total": len(files)} + is_image = fname.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")) + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": seed, + } + if is_image: + try: + with open(fpath, "rb") as f: + img_b64 = base64.b64encode(f.read()).decode() + body["images"] = [img_b64] + except Exception as e: + entry["error"] = f"Failed to read image: {e}" + entry["elapsed_ms"] = 0 + yield "data: " + json.dumps(entry) + "\n\n" + continue + else: + try: + with open(fpath, "r") as f: + body["contexts"] = [f.read()] + except Exception as e: + entry["error"] = f"Failed to read text file: {e}" + entry["elapsed_ms"] = 0 + yield "data: " + json.dumps(entry) + "\n\n" + continue + + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + elapsed = time.time() - t0 + entry["elapsed_ms"] = round(elapsed * 1000, 1) + if resp.status_code == 200: + entry["result"] = resp.json() + entry["status"] = "ok" + else: + entry["error"] = f"HTTP {resp.status_code}: {resp.text[:200]}" + entry["status"] = "error" + except Exception as e: + entry["error"] = str(e) + entry["elapsed_ms"] = round((time.time() - t0) * 1000, 1) + entry["status"] = "error" + + # Include image data for display if this is an image file + if is_image: + entry["is_image"] = True + with open(fpath, "rb") as f: + entry["image_data"] = base64.b64encode(f.read()).decode() + else: + entry["is_image"] = False + + yield "data: " + json.dumps(entry) + "\n\n" + + total_ms = round((time.time() - t_start) * 1000, 1) + yield "data: " + json.dumps({"event": "done", "total_files": len(files), "total_ms": total_ms}) + "\n\n" + + return Response(event_stream(), mimetype="text/event-stream") + + +@app.route("/api/image-selection/stream", methods=["POST"]) +def image_selection_stream(): + """Scan a folder of images and select which ONE image matches a question. + + Sends ALL images in a SINGLE /decision request as parallel contexts + (one context per image). The schema uses binary yes/no choices, so the + model evaluates each image against the question. Results are aggregated + to identify the single best match — only the winning image gets a + probability shown, satisfying the constraint that the model picks ONE + image from all options presented together. + + This approach is used because Qwen2.5-VL's vision architecture cannot + reliably associate text labels with specific images when multiple images + are interleaved in a single multimodal context. The yes/no approach + evaluates each image in its own dedicated context while all images are + still processed together in one batched decision pass. + + SSE events: + start: total_images, image_files, question + progress: per-image file loaded + result: per-image processing update + decision: single winning image with file name and probability + done: total_ms + """ + data = request.get_json(force=True, silent=True) or {} + source = data.get("source", "") + question = data.get("question", "Which image contains a yellow rubber duck?") + instructions = data.get("instructions", + "Look at this image and determine if it contains the described object. Answer YES if it does, NO if it does not.") + seed = data.get("seed", 42) + + # Gather image files (only images, sorted) + files = [] + if os.path.isdir(source): + for fname in sorted(os.listdir(source)): + fpath = os.path.join(source, fname) + if not os.path.isfile(fpath): + continue + lower = fname.lower() + if lower.endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")): + files.append((fpath, fname)) + elif os.path.isfile(source) and source.lower().endswith((".jpg", ".jpeg", ".png", ".bmp", ".webp", ".gif")): + files = [(source, os.path.basename(source))] + + def event_stream(): + if not files: + yield "data: " + json.dumps({"error": "No image files found", "total_images": 0}) + "\n\n" + return + + image_names = [fname for (_, fname) in files] + + # All images processed in a single decision request with yes/no schema. + # Each image is a separate context so the model can clearly associate + # the question with that one image. + schema = {"properties": {"match": {"type": "string", "enum": ["yes", "no"]}}} + contexts = [question for _ in files] + + images_b64 = [] + for idx, (fpath, fname) in enumerate(files): + try: + with open(fpath, "rb") as f: + images_b64.append(base64.b64encode(f.read()).decode()) + yield "data: " + json.dumps({"event": "progress", "file": fname, "loaded": True, "index": idx, "total": len(files)}) + "\n\n" + except Exception as e: + yield "data: " + json.dumps({"error": f"Failed to read {fname}: {e}", "file": fname}) + "\n\n" + return + + yield "data: " + json.dumps({ + "event": "start", + "total_images": len(files), + "image_files": image_names, + "question": question + }) + "\n\n" + + body = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "images": images_b64, + "seed": seed, + } + + t0 = time.time() + elapsed = 0.0 + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=300) + elapsed = time.time() - t0 + + if resp.status_code == 200: + data = resp.json() + results = data.get("results", []) + + # Aggregate: find images with "yes" answer, sorted by confidence + yes_matches = [] + for i, res in enumerate(results): + field = res.get("fields", {}).get("match", {}) + value = field.get("value", "no") + probability = field.get("probability", 0.0) + tokens = res.get("usage", {}).get("context_tokens", 0) + + entry = { + "file": image_names[i], + "match": value == "yes", + "probability": probability, + "value": value, + "context_tokens": tokens, + "index": i, + "total": len(results), + "is_image": True, + } + yield "data: " + json.dumps({"event": "result", "data": entry}) + "\n\n" + + if value == "yes": + yes_matches.append((i, probability, image_names[i])) + + # Select the single best match (highest yes confidence) + # Only one image gets reported as the decision winner + if yes_matches: + yes_matches.sort(key=lambda x: -x[1]) # Sort by probability descending + best = yes_matches[0] + idx, prob, fname = best + + decision = { + "selected_file": fname, + "selected_index": idx, + "probability": prob, + "question": question, + "total_matches": len(yes_matches), + "all_matches": [{"file": f, "probability": p} for (_, p, f) in yes_matches], + "image_data": images_b64[idx] if 0 <= idx < len(images_b64) else None, + "is_image": True, + "timings": data.get("timings", {}), + } + yield "data: " + json.dumps({"event": "decision", "data": decision}) + "\n\n" + else: + # No matches found — report the highest "no" confidence as near-miss + all_probs = [] + for i, res in enumerate(results): + field = res.get("fields", {}).get("match", {}) + prob = field.get("probability", 0.0) + all_probs.append((i, prob, image_names[i])) + + # Find the one with highest "yes" probability even if it answered "no" + # This gives useful info about confidence + decision = { + "selected_file": None, + "selected_index": -1, + "probability": 0.0, + "question": question, + "total_matches": 0, + "timings": data.get("timings", {}), + } + yield "data: " + json.dumps({"event": "decision", "data": decision}) + "\n\n" + + else: + yield "data: " + json.dumps({"error": f"HTTP {resp.status_code}: {resp.text[:300]}"}) + "\n\n" + except Exception as e: + yield "data: " + json.dumps({"error": str(e)}) + "\n\n" + + yield "data: " + json.dumps({"event": "done", "total_ms": round(elapsed * 1000, 1)}) + "\n\n" + + return Response(event_stream(), mimetype="text/event-stream") + + +@app.route("/api/export-snippet", methods=["POST"]) +def export_snippet(): + data = request.get_json(force=True, silent=True) or {} + schema_json = json.dumps(data.get("schema", {}), indent=2) if data.get("schema") else "{}" + contexts_json = json.dumps(data.get("contexts", []), indent=2) + instructions = data.get("instructions", "") + + snippet = '''"""Vision Decision API request snippet. + +Endpoint: POST ''' + SERVER_URL + '''/decision + +Request body: + contexts: list[str] -- one string per decision context + schema : dict -- JSON Schema with "properties" (required) where each + property defines an enum/integer/number/boolean field + images : list[str] -- optional, base64-encoded image data (one per media marker in context) + instructions : str -- appended to the system prompt + seed : int -- optional, for reproducibility + +Response format: + { + "object": "decision", + "results": [ + { + "decision": { "action": "CLICK_OK" }, + "fields": { "action": { "value": "CLICK_OK", "probability": 0.99, "scored_nodes": 2, "tree": true }}, + "usage": { "context_tokens": 18, "scored_rows": 11 } + } + ], + "timings": { + "prefill_ms": 340.5, + "scoring_ms": 12.3, + "total_ms": 352.8, + "rounds": 1 + } + } +""" + +import base64, requests, json + +schema = ''' + schema_json + ''' +contexts = ''' + contexts_json + ''' +instructions = ''' + json.dumps(instructions) + ''' + +payload = { + "contexts": contexts, + "schema": schema, + "instructions": instructions, + "seed": 42, +} + +# Optional: add images to payload if needed: +# with open("path/to/image.jpg", "rb") as f: +# img_b64 = base64.b64encode(f.read()).decode() +# payload["images"] = [img_b64] + +resp = requests.post("''' + SERVER_URL + '''/decision", json=payload, timeout=120) +print(json.dumps(resp.json(), indent=2)) +''' + return jsonify({"snippet": snippet}) + + +if __name__ == "__main__": + port = int(os.environ.get("PORT", 5786)) + app.run(host="0.0.0.0", port=port, debug=False) + diff --git a/tools/parallel-decision/examples/vision-decision-harness/requirements.txt b/tools/parallel-decision/examples/vision-decision-harness/requirements.txt new file mode 100644 index 000000000000..690a3cefbe89 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/requirements.txt @@ -0,0 +1,3 @@ +flask +flask-cors +requests diff --git a/tools/parallel-decision/examples/vision-decision-harness/templates/index.html b/tools/parallel-decision/examples/vision-decision-harness/templates/index.html new file mode 100644 index 000000000000..be2256c9b6da --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/templates/index.html @@ -0,0 +1,749 @@ + + + + + + Vision Decision Harness + + + +
+

Vision Decision Harness

+
Harness on 0.0.0.0:5786 | Proxies to llama-server at 0.0.0.0:8081
+ +
+ +
+
+

Source

+
+ +
+ + +
+ + Drag-drop or browse a folder of images (.jpg/.png/etc) and text files (.txt/.md) +
+
+ + +
+
+ +
+

Prompt & Tags

+
+ + +
+
+
+ + +
+
+ + +
+
+ + +
+
+ +
+ + + +
+
+
+
+
+ + +
+
+ +
+

Actions

+
+ +
+
+ +
+
+ +
+
+ +
+
+ +
+
+
+ + +
+
+

Results

+
+
+
+ + + + + + +
FileImageDecisionProbTokensTime (ms)Status
+ + +
+

Log

+
+
+ + + + + + + diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json b/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json new file mode 100644 index 000000000000..f002a3165b36 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/answer_key.json @@ -0,0 +1,190 @@ +{ + "description": "Answer key for decision scoring test suite", + "total_tests": 23, + "tests": [ + { + "id": "t01_basic_color", + "context": "What is the dominant color in this image?", + "expected": "brown", + "difficulty": "easy", + "category": "color", + "schema_enum": ["red", "green", "blue", "brown", "orange"] + }, + { + "id": "t02_animal_type", + "context": "What type of animal is in this image?", + "expected": "dog", + "difficulty": "medium", + "category": "animal", + "schema_enum": ["dog", "cat", "bird", "horse", "fish"] + }, + { + "id": "t03_ui_element", + "context": "What type of UI element should you click to submit this form?", + "expected": "button", + "difficulty": "easy", + "category": "ui", + "schema_enum": ["button", "link", "checkbox", "text_field", "dropdown"] + }, + { + "id": "t04_action_select", + "context": "Based on this content, what is the best action to take?", + "expected": "CLICK_OK", + "difficulty": "easy", + "category": "action", + "schema_enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"] + }, + { + "id": "t05_sentiment", + "context": "What is the sentiment expressed in this content?", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "schema_enum": ["positive", "negative", "neutral"] + }, + { + "id": "t06_navigation", + "context": "Where should you navigate to complete this task?", + "expected": "settings", + "difficulty": "hard", + "category": "navigation", + "schema_enum": ["home", "settings", "profile", "logout", "help"] + }, + { + "id": "t07_content_type", + "context": "What type of content is shown in this image?", + "expected": "image", + "difficulty": "easy", + "category": "content", + "schema_enum": ["text", "image", "video", "audio", "interactive"] + }, + { + "id": "t08_priority", + "context": "What priority level should this task be assigned?", + "expected": "medium", + "difficulty": "hard", + "category": "priority", + "schema_enum": ["low", "medium", "high", "urgent"] + }, + { + "id": "t09_safety", + "context": "Is it safe to proceed with this action?", + "expected": "yes", + "difficulty": "medium", + "category": "safety", + "schema_enum": ["yes", "no", "caution"] + }, + { + "id": "t10_layout", + "context": "Where should new items be added in this layout?", + "expected": "bottom_right", + "difficulty": "hard", + "category": "layout", + "schema_enum": ["top_left", "top_right", "bottom_left", "bottom_right", "center"] + }, + { + "id": "t11_text_format", + "context": "What format is this text in?", + "expected": "plain", + "difficulty": "medium", + "category": "format", + "schema_enum": ["plain", "markdown", "html", "json", "xml"] + }, + { + "id": "t12_object_count", + "context": "How many main objects are visible in this image?", + "expected": "one", + "difficulty": "medium", + "category": "counting", + "schema_enum": ["one", "two", "three", "four", "many"] + }, + { + "id": "t13_file_op", + "context": "What file operation should be performed?", + "expected": "save", + "difficulty": "easy", + "category": "fileop", + "schema_enum": ["save", "delete", "rename", "copy", "move"] + }, + { + "id": "t14_error_type", + "context": "What type of error is indicated by this message?", + "expected": "not_found", + "difficulty": "hard", + "category": "error", + "schema_enum": ["syntax", "runtime", "network", "permission", "not_found"] + }, + { + "id": "t15_time_urgency", + "context": "How urgent is this action?", + "expected": "soon", + "difficulty": "hard", + "category": "time", + "schema_enum": ["immediate", "soon", "later", "anytime"] + }, + { + "id": "t16_language", + "context": "What language is this text written in?", + "expected": "english", + "difficulty": "hard", + "category": "language", + "schema_enum": ["english", "spanish", "french", "german", "chinese"] + }, + { + "id": "t17_component", + "context": "What type of UI component is shown here?", + "expected": "card", + "difficulty": "medium", + "category": "ui", + "schema_enum": ["modal", "sidebar", "header", "footer", "card"] + }, + { + "id": "t18_data_source", + "context": "Where should this data be fetched from?", + "expected": "api", + "difficulty": "hard", + "category": "data", + "schema_enum": ["database", "api", "file", "cache", "user_input"] + }, + { + "id": "t19_interaction", + "context": "What type of user interaction does this represent?", + "expected": "click", + "difficulty": "medium", + "category": "interaction", + "schema_enum": ["click", "hover", "drag", "scroll", "type"] + }, + { + "id": "t20_state", + "context": "What is the correct state for this toggle?", + "expected": "on", + "difficulty": "hard", + "category": "state", + "schema_enum": ["on", "off", "indeterminate", "disabled"] + }, + { + "id": "t21_direction", + "context": "Which direction should this carousel move?", + "expected": "right", + "difficulty": "hard", + "category": "direction", + "schema_enum": ["left", "right", "up", "down", "none"] + }, + { + "id": "t22_file_type", + "context": "What type of file is this?", + "expected": "image", + "difficulty": "easy", + "category": "filetype", + "schema_enum": ["image", "document", "spreadsheet", "presentation", "archive"] + }, + { + "id": "t23_security", + "context": "What security level is required for this resource?", + "expected": "internal", + "difficulty": "hard", + "category": "security", + "schema_enum": ["public", "internal", "confidential", "restricted"] + } + ] +} diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py b/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py new file mode 100644 index 000000000000..98ef30836fdb --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/evaluate_text_tests.py @@ -0,0 +1,195 @@ +#!/usr/bin/env python3 +""" +Evaluation runner for the text-based decision test suite. + +Processes all .txt files in tests/text_samples/ through the llama-server +/decision endpoint using the shared sentiment schema, then compares results +against answer_key.json. + +Usage: + python3 evaluate_text_tests.py + +Environment: + LLAMA_SERVER_URL - defaults to http://0.0.0.0:8081 +""" + +import json +import os +import sys +import time +import requests + +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") +SAMPLES_DIR = os.path.join(os.path.dirname(__file__), "text_samples") + + +def load_answer_key(): + path = os.path.join(SAMPLES_DIR, "answer_key.json") + with open(path) as f: + return json.load(f) + + +def run_single_decision(context_text, schema, instructions, seed=42): + """Send a single decision request to the llama-server.""" + body = { + "contexts": [context_text], + "schema": schema, + "instructions": instructions, + "seed": seed, + } + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=120) + return resp + + +def main(): + key_data = load_answer_key() + schema = key_data["schema"] + question = key_data["question"] + tests = key_data["tests"] + expected_field = list(schema["properties"].keys())[0] # "sentiment" + + print(f"Text Decision Test Suite") + print(f"Server: {SERVER_URL}") + print(f"Question: {question}") + print(f"Schema: {json.dumps(schema['properties'])}") + print(f"Tests: {len(tests)}") + print() + + # Check server is reachable + try: + requests.get(f"{SERVER_URL}/v1/models", timeout=5) + except Exception: + print(f"ERROR: Cannot reach llama-server at {SERVER_URL}") + print("The server must be running for tests to execute.") + return 1 + + instructions = "Classify the sentiment expressed in the following text." + results = [] + t_total = time.time() + + for i, test in enumerate(tests): + filepath = os.path.join(SAMPLES_DIR, test["file"]) + try: + with open(filepath, "r") as f: + text = f.read().strip() + except Exception as e: + results.append({"file": test["file"], "error": str(e), "correct": False}) + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} ERROR: {e}") + continue + + t0 = time.time() + resp = run_single_decision(text, schema, instructions) + elapsed = time.time() - t0 + + if resp.status_code == 200: + data = resp.json() + result = data.get("results", [{}])[0] + fields = result.get("fields", {}) + field_data = fields.get(expected_field, {}) + predicted = field_data.get("value", "UNKNOWN") + probability = field_data.get("probability", 0.0) + tokens = result.get("usage", {}).get("context_tokens", 0) + timings = data.get("timings", {}) + + expected = test["expected"] + correct = predicted == expected + status = "PASS" if correct else "FAIL" + prob_str = f"{probability*100:.1f}%" + + mark = "" if correct else f" [expected: {expected}]" + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} {status} -> {predicted:<10} ({prob_str}) [{elapsed*1000:.0f}ms, {tokens} toks]{mark}") + + results.append({ + "file": test["file"], + "predicted": predicted, + "expected": expected, + "correct": correct, + "probability": probability, + "elapsed_ms": round(elapsed * 1000, 1), + "context_tokens": tokens, + "difficulty": test["difficulty"], + "category": test["category"], + "timings": timings, + }) + else: + print(f"[{i+1:2d}/{len(tests)}] {test['file']:<8} HTTP {resp.status_code}: {resp.text[:100]}") + results.append({ + "file": test["file"], + "error": f"HTTP {resp.status_code}", + "correct": False, + "difficulty": test["difficulty"], + "category": test["category"], + }) + + total_elapsed = time.time() - t_total + + # Print summary + print() + print("=" * 80) + print("SUMMARY") + print("=" * 80) + + correct_count = sum(1 for r in results if r.get("correct")) + total_count = len(results) + print(f"Accuracy: {correct_count}/{total_count} ({correct_count/total_count*100:.1f}%)") + print(f"Total time: {total_elapsed:.1f}s ({total_elapsed/total_count*1000:.0f}ms per test avg)") + + # By difficulty + print("\nBy Difficulty:") + for diff in ["easy", "medium", "hard"]: + diff_results = [r for r in results if r.get("difficulty") == diff] + if diff_results: + diff_correct = sum(1 for r in diff_results if r.get("correct")) + correct_probs = [r.get("probability", 0) for r in diff_results if r.get("correct")] + wrong_probs = [r.get("probability", 0) for r in diff_results if not r.get("correct") and "probability" in r] + avg_prob = sum(correct_probs) / len(correct_probs) if correct_probs else 0 + avg_wrong = sum(wrong_probs) / len(wrong_probs) if wrong_probs else 0 + print(f" {diff:<8}: {diff_correct}/{len(diff_results)} ({diff_correct/len(diff_results)*100:.0f}%) avg_prob={avg_prob:.1%} avg_wrong_prob={avg_wrong:.1%}") + + # By category + print("\nBy Category:") + categories = sorted(set(r.get("category", "") for r in results)) + for cat in categories: + cat_results = [r for r in results if r.get("category") == cat] + cat_correct = sum(1 for r in cat_results if r.get("correct")) + print(f" {cat:<14}: {cat_correct}/{len(cat_results)} ({cat_correct/len(cat_results)*100:.0f}%)") + + # Calibration check + correct_probs = [r["probability"] for r in results if r.get("correct") and "probability" in r] + wrong_probs = [r["probability"] for r in results if not r.get("correct") and "probability" in r] + if correct_probs and wrong_probs: + avg_correct = sum(correct_probs) / len(correct_probs) + avg_wrong = sum(wrong_probs) / len(wrong_probs) + print(f"\nCalibration:") + print(f" Avg confidence (correct answers): {avg_correct:.1%}") + print(f" Avg confidence (wrong answers): {avg_wrong:.1%}") + print(f" Well calibrated: {'YES' if avg_correct > avg_wrong else 'NO'}") + + # Wrong answers + wrong = [r for r in results if not r.get("correct")] + if wrong: + print(f"\nWrong answers ({len(wrong)}):") + for r in wrong: + print(f" {r['file']}: expected '{r.get('expected', '?')}', got '{r.get('predicted', '?')}' ({r.get('probability', 0):.1%})") + + # Save results + results_path = os.path.join(SAMPLES_DIR, "evaluation_results.json") + with open(results_path, "w") as f: + json.dump({ + "question": question, + "schema": schema, + "total_tests": total_count, + "passed": correct_count, + "failed": total_count - correct_count, + "accuracy": f"{correct_count/total_count*100:.1f}%", + "total_time_ms": round(total_elapsed * 1000, 1), + "avg_time_ms": round(total_elapsed / total_count * 1000, 1), + "results": results, + }, f, indent=2) + print(f"\nDetailed results saved to {results_path}") + + return 0 if correct_count == total_count else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py b/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py new file mode 100644 index 000000000000..ccc7141de4cc --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/test_decision_scoring.py @@ -0,0 +1,633 @@ +#!/usr/bin/env python3 +""" +Test suite for evaluating the llama-server /decision endpoint's +scoring and classification abilities. + +These are TEXT-ONLY tests — no image references. Each test presents +a text context and asks the model to classify or choose among options +based on the textual content alone. +""" + +import json +import os +import base64 +import time +import requests +from pathlib import Path + +# ---- Configuration ---- +SERVER_URL = os.environ.get("LLAMA_SERVER_URL", "http://0.0.0.0:8081") +TEST_IMAGE_DIR = os.environ.get("TEST_IMAGE_DIR", + os.path.join(os.path.dirname(os.path.dirname(__file__)), "..", "llama.cpp-thecodacus", "tools", "mtmd")) + + +# ---- Test Cases (TEXT-ONLY, no image references) ---- +TESTS = [ + # 1. Sentiment classification + { + "id": "t01_sentiment_pos", + "context": "The product arrived early and exceeded all expectations. The packaging was perfect and the quality is outstanding.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + }, + # 2. Sentiment - negative + { + "id": "t02_sentiment_neg", + "context": "This service was terrible. The staff was rude, the food was cold, and I had to wait an hour. Never coming back.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + }, + # 3. Sentiment - neutral + { + "id": "t03_sentiment_neu", + "context": "The meeting is scheduled for 3 PM in conference room B. Please bring the quarterly report.", + "schema": {"properties": {"sentiment": {"type": "string", + "enum": ["positive", "negative", "neutral"]}}}, + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + }, + # 4. Urgency classification + { + "id": "t04_urgency_high", + "context": "The production server is down and customers cannot complete purchases. Immediate action required.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "critical", + "difficulty": "easy", + "category": "urgency", + }, + # 5. Urgency - medium + { + "id": "t05_urgency_med", + "context": "Please review the design documents and provide feedback by end of week.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "medium", + "difficulty": "medium", + "category": "urgency", + }, + # 6. Urgency - low + { + "id": "t06_urgency_low", + "context": "When you have time, could you update the project wiki with the new API endpoints?", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["low", "medium", "high", "critical"]}}}, + "expected": "low", + "difficulty": "hard", + "category": "urgency", + }, + # 7. Action selection + { + "id": "t07_action_ok", + "context": "The user has confirmed their email address and completed the registration form. All validation checks passed.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "CLICK_OK", + "difficulty": "easy", + "category": "action", + }, + # 8. Action - cancel + { + "id": "t08_action_cancel", + "context": "The payment was declined due to insufficient funds. The user needs to provide alternative payment.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "CLICK_CANCEL", + "difficulty": "medium", + "category": "action", + }, + # 9. Action - no action + { + "id": "t09_action_none", + "context": "All system checks passed. The service is running normally with no issues detected. No further action needed.", + "schema": {"properties": {"action": {"type": "string", + "enum": ["CLICK_OK", "CLICK_CANCEL", "NO_ACTION"]}}}, + "expected": "NO_ACTION", + "difficulty": "hard", + "category": "action", + }, + # 10. Content type classification + { + "id": "t10_content_news", + "context": "Breaking: The city council voted today to approve the new budget proposal. The measure passed with a 7-2 majority after three hours of debate.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "news", + "difficulty": "easy", + "category": "content", + }, + # 11. Content type - instruction + { + "id": "t11_content_instruction", + "context": "To reset your password, first navigate to the login page. Click the 'Forgot Password' link. Enter your email address and submit.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "instruction", + "difficulty": "medium", + "category": "content", + }, + # 12. Content type - summary + { + "id": "t12_content_summary", + "context": "In summary, the experiment demonstrated a significant improvement in performance. Key findings include a 15% increase in throughput and 20% reduction in latency compared to baseline.", + "schema": {"properties": {"type": {"type": "string", + "enum": ["news", "opinion", "advertisement", "instruction", "summary"]}}}, + "expected": "summary", + "difficulty": "medium", + "category": "content", + }, + # 13. Priority classification + { + "id": "t13_priority_high", + "context": "Security vulnerability detected in production. SQL injection risk in the user authentication endpoint. Fix immediately.", + "schema": {"properties": {"priority": {"type": "string", + "enum": ["low", "medium", "high", "urgent"]}}}, + "expected": "urgent", + "difficulty": "medium", + "category": "priority", + }, + # 14. Priority - low + { + "id": "t14_priority_low", + "context": "Consider updating the style guide documentation when time permits next quarter.", + "schema": {"properties": {"priority": {"type": "string", + "enum": ["low", "medium", "high", "urgent"]}}}, + "expected": "low", + "difficulty": "hard", + "category": "priority", + }, + # 15. File operation + { + "id": "t15_file_save", + "context": "The user has finished editing the document and clicked the save button. The changes should be persisted to disk.", + "schema": {"properties": {"operation": {"type": "string", + "enum": ["save", "delete", "rename", "copy", "move"]}}}, + "expected": "save", + "difficulty": "easy", + "category": "fileop", + }, + # 16. File operation - delete + { + "id": "t16_file_delete", + "context": "The temporary cache files from last week's processing run are no longer needed and should be removed to free up space.", + "schema": {"properties": {"operation": {"type": "string", + "enum": ["save", "delete", "rename", "copy", "move"]}}}, + "expected": "delete", + "difficulty": "medium", + "category": "fileop", + }, + # 17. File type + { + "id": "t17_filetype_image", + "context": "File: photo_2024_09_25_143022.jpg — JPEG image, 1920x1080 pixels, ICC profile: sRGB, EXIF data present, camera: iPhone 14 Pro.", + "schema": {"properties": {"filetype": {"type": "string", + "enum": ["image", "document", "spreadsheet", "presentation", "archive"]}}}, + "expected": "image", + "difficulty": "easy", + "category": "filetype", + }, + # 18. File type - document + { + "id": "t18_filetype_doc", + "context": "File: quarterly_report.pdf — PDF document, 42 pages, contains tables, charts, and formatted text. Last modified: 2024-09-20.", + "schema": {"properties": {"filetype": {"type": "string", + "enum": ["image", "document", "spreadsheet", "presentation", "archive"]}}}, + "expected": "document", + "difficulty": "medium", + "category": "filetype", + }, + # 19. Error type + { + "id": "t19_error_notfound", + "context": "Error: Resource not found at /api/users/12345. The requested user ID does not exist in the database.", + "schema": {"properties": {"error": {"type": "string", + "enum": ["syntax", "runtime", "network", "permission", "not_found"]}}}, + "expected": "not_found", + "difficulty": "hard", + "category": "error", + }, + # 20. Error type - permission + { + "id": "t20_error_permission", + "context": "Access denied: User does not have permission to read /etc/shadow. Required role: root, current role: standard_user.", + "schema": {"properties": {"error": {"type": "string", + "enum": ["syntax", "runtime", "network", "permission", "not_found"]}}}, + "expected": "permission", + "difficulty": "hard", + "category": "error", + }, + # 21. UI element type + { + "id": "t21_ui_button", + "context": "The form contains a blue rectangular button labeled 'Submit Order'. Clicking it sends the form data to the server.", + "schema": {"properties": {"element": {"type": "string", + "enum": ["button", "link", "checkbox", "text_field", "dropdown"]}}}, + "expected": "button", + "difficulty": "easy", + "category": "ui", + }, + # 22. UI element - dropdown + { + "id": "t22_ui_dropdown", + "context": "The settings panel shows a dropdown menu with options: Light Mode, Dark Mode, Auto. The current selection is Dark Mode.", + "schema": {"properties": {"element": {"type": "string", + "enum": ["button", "link", "checkbox", "text_field", "dropdown"]}}}, + "expected": "dropdown", + "difficulty": "medium", + "category": "ui", + }, + # 23. Language detection + { + "id": "t23_lang_english", + "context": "The quick brown fox jumps over the lazy dog. Pack my box with five dozen liquor jugs. How vexingly quick daft zebras jump!", + "schema": {"properties": {"lang": {"type": "string", + "enum": ["english", "spanish", "french", "german", "chinese"]}}}, + "expected": "english", + "difficulty": "easy", + "category": "language", + }, + # 24. Language - Spanish + { + "id": "t24_lang_spanish", + "context": "Hola, como estas? Me encantaria visitar Espana algun dia. La comida espanola es deliciosa, especialmente la paella.", + "schema": {"properties": {"lang": {"type": "string", + "enum": ["english", "spanish", "french", "german", "chinese"]}}}, + "expected": "spanish", + "difficulty": "medium", + "category": "language", + }, + # 25. Navigation + { + "id": "t25_navigation_settings", + "context": "To change your notification preferences, go to the account section and look for the bell icon. Click it to open the settings panel.", + "schema": {"properties": {"destination": {"type": "string", + "enum": ["home", "settings", "profile", "logout", "help"]}}}, + "expected": "settings", + "difficulty": "hard", + "category": "navigation", + }, + # 26. Format detection + { + "id": "t26_format_json", + "context": '{"users": [{"name": "Alice", "id": 1}, {"name": "Bob", "id": 2}], "count": 2}', + "schema": {"properties": {"format": {"type": "string", + "enum": ["plain", "markdown", "html", "json", "xml"]}}}, + "expected": "json", + "difficulty": "easy", + "category": "format", + }, + # 27. Format - markdown + { + "id": "t27_format_markdown", + "context": "# Project README\n\nThis project does **important things**. See the [documentation](docs.md) for details.\n\n```python\nprint('hello')\n```", + "schema": {"properties": {"format": {"type": "string", + "enum": ["plain", "markdown", "html", "json", "xml"]}}}, + "expected": "markdown", + "difficulty": "medium", + "category": "format", + }, + # 28. Interaction type + { + "id": "t28_interaction_click", + "context": "The user pressed the red submit button with their mouse. The button highlighted blue briefly and then the form was submitted.", + "schema": {"properties": {"interaction": {"type": "string", + "enum": ["click", "hover", "drag", "scroll", "type"]}}}, + "expected": "click", + "difficulty": "easy", + "category": "interaction", + }, + # 29. Interaction - scroll + { + "id": "t29_interaction_scroll", + "context": "The user moved the scrollbar down using the mouse wheel to see more content below the fold of the webpage.", + "schema": {"properties": {"interaction": {"type": "string", + "enum": ["click", "hover", "drag", "scroll", "type"]}}}, + "expected": "scroll", + "difficulty": "hard", + "category": "interaction", + }, + # 30. State classification + { + "id": "t30_state_on", + "context": "The switch is in the active position. The LED indicator is lit green. Power is flowing to the connected device.", + "schema": {"properties": {"state": {"type": "string", + "enum": ["on", "off", "indeterminate", "disabled"]}}}, + "expected": "on", + "difficulty": "easy", + "category": "state", + }, + # 31. State - off + { + "id": "t31_state_off", + "context": "The device is powered down. The LED is unlit. No power is being drawn from the battery. The switch is in the inactive position.", + "schema": {"properties": {"state": {"type": "string", + "enum": ["on", "off", "indeterminate", "disabled"]}}}, + "expected": "off", + "difficulty": "medium", + "category": "state", + }, + # 32. Data source + { + "id": "t32_data_api", + "context": "The frontend fetches user data by making a GET request to /api/v2/users with authentication headers. Results are returned as JSON.", + "schema": {"properties": {"source": {"type": "string", + "enum": ["database", "api", "file", "cache", "user_input"]}}}, + "expected": "api", + "difficulty": "medium", + "category": "data", + }, + # 33. Data source - user_input + { + "id": "t33_data_user", + "context": "The search query was typed directly into the search box by the user. No predefined data source was queried.", + "schema": {"properties": {"source": {"type": "string", + "enum": ["database", "api", "file", "cache", "user_input"]}}}, + "expected": "user_input", + "difficulty": "hard", + "category": "data", + }, + # 34. Security level + { + "id": "t34_security_internal", + "context": "This document contains company-wide policies for internal use only. It is accessible to all employees but not to external parties.", + "schema": {"properties": {"level": {"type": "string", + "enum": ["public", "internal", "confidential", "restricted"]}}}, + "expected": "internal", + "difficulty": "medium", + "category": "security", + }, + # 35. Security - public + { + "id": "t35_security_public", + "context": "The marketing brochure is available on the company website for anyone to download and share freely.", + "schema": {"properties": {"level": {"type": "string", + "enum": ["public", "internal", "confidential", "restricted"]}}}, + "expected": "public", + "difficulty": "easy", + "category": "security", + }, + # 36. Time urgency + { + "id": "t36_time_soon", + "context": "Please review the draft proposal and provide feedback within the next few days. The deadline is Friday.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["immediate", "soon", "later", "anytime"]}}}, + "expected": "soon", + "difficulty": "hard", + "category": "time", + }, + # 37. Time urgency - later + { + "id": "t37_time_later", + "context": "When you have time over the next few weeks, please update the project documentation with the new API changes.", + "schema": {"properties": {"urgency": {"type": "string", + "enum": ["immediate", "soon", "later", "anytime"]}}}, + "expected": "later", + "difficulty": "hard", + "category": "time", + }, + # 38. Direction + { + "id": "t38_direction_right", + "context": "The carousel should advance to the next item. The user clicked the right arrow button to move forward.", + "schema": {"properties": {"direction": {"type": "string", + "enum": ["left", "right", "up", "down", "none"]}}}, + "expected": "right", + "difficulty": "medium", + "category": "direction", + }, + # 39. Direction - left + { + "id": "t39_direction_left", + "context": "The user wants to go back to the previous slide. Click the left arrow to navigate backwards.", + "schema": {"properties": {"direction": {"type": "string", + "enum": ["left", "right", "up", "down", "none"]}}}, + "expected": "left", + "difficulty": "hard", + "category": "direction", + }, + # 40. Component type + { + "id": "t40_component_card", + "context": "Each user profile is displayed in a bordered box with a shadow. It contains the profile picture, name, and status below.", + "schema": {"properties": {"component": {"type": "string", + "enum": ["modal", "sidebar", "header", "footer", "card"]}}}, + "expected": "card", + "difficulty": "medium", + "category": "ui", + }, + # 41. Component - sidebar + { + "id": "t41_component_sidebar", + "context": "The navigation panel on the left side of the screen contains links to Dashboard, Settings, and Logout.", + "schema": {"properties": {"component": {"type": "string", + "enum": ["modal", "sidebar", "header", "footer", "card"]}}}, + "expected": "sidebar", + "difficulty": "hard", + "category": "ui", + }, +] + + +def run_single_decision(test_case, timeout=120): + """Run a single decision test and return the result dict.""" + body = { + "contexts": [test_case["context"]], + "schema": test_case["schema"], + "instructions": "Select the correct value from the allowed choices based on the context provided.", + "seed": 42, + } + + t0 = time.time() + try: + resp = requests.post(f"{SERVER_URL}/decision", json=body, timeout=timeout) + elapsed = time.time() - t0 + if resp.status_code == 200: + data = resp.json() + field_name = list(data["results"][0]["fields"].keys())[0] + fld = data["results"][0]["fields"][field_name] + return { + "success": True, + "decision": fld["value"], + "probability": fld["probability"], + "expected": test_case["expected"], + "elapsed_ms": round(elapsed * 1000), + "tokens": data["results"][0]["usage"].get("context_tokens", 0), + "timings": data.get("timings", {}), + "id": test_case["id"], + "context": test_case["context"], + "difficulty": test_case["difficulty"], + "category": test_case["category"], + } + else: + return { + "success": False, + "error": f"HTTP {resp.status_code}: {resp.text[:200]}", + "id": test_case["id"], + } + except Exception as e: + return { + "success": False, + "error": str(e), + "id": test_case["id"], + } + + +def evaluate_results(results): + """Evaluate test results and print a summary.""" + passed = 0 + failed = 0 + total_prob_correct = 0.0 + total_prob_wrong = 0.0 + correct_count = 0 + wrong_count = 0 + + print("\n" + "=" * 90) + print("DECISION SCORING TEST RESULTS (TEXT-ONLY)") + print("=" * 90) + print(f"{'ID':<18} {'Diff':<8} {'Category':<14} {'Expected':<14} {'Got':<14} {'Prob':<8} {'Time':<8} {'Status':<6}") + print("-" * 90) + + by_difficulty = {} + by_category = {} + + for r in results: + if not r["success"]: + print(f"{r['id']:<18} {'ERROR':<8} {'':<14} {'N/A':<14} {'N/A':<14} {'N/A':<8} {'N/A':<8} FAIL") + failed += 1 + continue + + is_correct = r["decision"] == r["expected"] + status = "PASS" if is_correct else "FAIL" + if is_correct: + passed += 1 + total_prob_correct += r["probability"] + correct_count += 1 + else: + failed += 1 + total_prob_wrong += r["probability"] + wrong_count += 1 + + diff = r["difficulty"] + if diff not in by_difficulty: + by_difficulty[diff] = {"passed": 0, "total": 0, "probs": []} + by_difficulty[diff]["total"] += 1 + if is_correct: + by_difficulty[diff]["passed"] += 1 + by_difficulty[diff]["probs"].append(r["probability"]) + + cat = r["category"] + if cat not in by_category: + by_category[cat] = {"passed": 0, "total": 0, "probs": []} + by_category[cat]["total"] += 1 + if is_correct: + by_category[cat]["passed"] += 1 + by_category[cat]["probs"].append(r["probability"]) + + print(f"{r['id']:<18} {diff:<8} {cat:<14} {r['expected']:<14} {r['decision']:<14} {r['probability']:.2%} {str(r['elapsed_ms'])+'ms':<8} {status}") + + print("-" * 90) + print(f"\nOverall: {passed}/{len(results)} passed ({passed/len(results)*100:.1f}%)\n") + + print("By Difficulty:") + for diff in sorted(by_difficulty.keys()): + d = by_difficulty[diff] + avg_prob = sum(d["probs"]) / len(d["probs"]) if d["probs"] else 0 + print(f" {diff:<8}: {d['passed']}/{d['total']} ({d['passed']/d['total']*100:.1f}%) avg_prob={avg_prob:.2%}") + + print("\nBy Category:") + for cat in sorted(by_category.keys()): + c = by_category[cat] + avg_prob = sum(c["probs"]) / len(c["probs"]) if c["probs"] else 0 + print(f" {cat:<14}: {c['passed']}/{c['total']} ({c['passed']/c['total']*100:.1f}%) avg_prob={avg_prob:.2%}") + + # Calibration analysis + if correct_count > 0 and wrong_count > 0: + avg_correct = total_prob_correct / correct_count + avg_wrong = total_prob_wrong / wrong_count + print(f"\nCalibration:") + print(f" Avg prob (correct): {avg_correct:.2%}") + print(f" Avg prob (wrong): {avg_wrong:.2%}") + if avg_correct > avg_wrong: + print(f" -> WELL CALIBRATED (correct answers have higher confidence)") + else: + print(f" -> POORLY CALIBRATED (wrong answers have higher confidence)") + elif correct_count > 0: + print(f"\nCalibration: All answers correct ({total_prob_correct/correct_count:.2%} avg prob)") + + # Timing summary + valid_times = [r["elapsed_ms"] for r in results if r["success"]] + if valid_times: + print(f"\nTiming:") + print(f" Avg per decision: {sum(valid_times)/len(valid_times):.0f}ms") + print(f" Min: {min(valid_times)}ms | Max: {max(valid_times)}ms") + total_time = sum(valid_times) + print(f" Total: {total_time}ms ({total_time/1000:.1f}s)") + + return passed, failed + + +def save_answer_key(): + """Save the answer key separately from the test code.""" + key = { + "description": "Answer key for decision scoring test suite (text-only)", + "total_tests": len(TESTS), + "tests": [ + { + "id": t["id"], + "context": t["context"][:80] + "..." if len(t["context"]) > 80 else t["context"], + "expected": t["expected"], + "difficulty": t["difficulty"], + "category": t["category"], + "schema_enum": list(t["schema"]["properties"].values())[0]["enum"], + } + for t in TESTS + ], + } + + key_path = os.path.join(os.path.dirname(__file__), "answer_key.json") + with open(key_path, "w") as f: + json.dump(key, f, indent=2) + print(f"Answer key saved to {key_path}") + return key_path + + +if __name__ == "__main__": + import sys + + save_key = "--save-key" in sys.argv + + if save_key: + save_answer_key() + print("Answer key saved. Run without --save-key to test.") + sys.exit(0) + + print(f"Running {len(TESTS)} decision scoring tests (text-only)...") + print(f"Server: {SERVER_URL}") + + results = [] + for i, test in enumerate(TESTS): + r = run_single_decision(test) + results.append(r) + if r["success"]: + status = "PASS" if r["decision"] == r["expected"] else "FAIL" + print(f"[{i+1:2d}/{len(TESTS)}] {test['id']}: {status} -> {r['decision']} ({r['probability']:.1%}) [{r['elapsed_ms']}ms]") + else: + print(f"[{i+1:2d}/{len(TESTS)}] {test['id']}: ERROR: {r.get('error', 'unknown')}") + + passed, failed = evaluate_results(results) + + # Save results + results_path = os.path.join(os.path.dirname(__file__), "test_results.json") + with open(results_path, "w") as f: + json.dump(results, f, indent=2) + print(f"\nDetailed results saved to {results_path}") + + sys.exit(0 if failed == 0 else 1) diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json new file mode 100644 index 000000000000..83409546fd77 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/answer_key.json @@ -0,0 +1,159 @@ +{ + "description": "Answer key for text-based sentiment decision test suite. All 20 files use the same schema and question for consistent batch evaluation.", + "question": "What is the sentiment expressed in the following text?", + "schema": { + "properties": { + "sentiment": { + "type": "string", + "enum": [ + "positive", + "negative", + "neutral" + ] + } + } + }, + "total_files": 20, + "tests": [ + { + "file": "t01.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "This product is absolutely amazing! I love it and would recommend it to everyone" + }, + { + "file": "t02.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "This service was terrible. The staff was rude, the food was cold, and I had to w" + }, + { + "file": "t03.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The meeting is scheduled for 3 PM in conference room B. Please bring the quarter" + }, + { + "file": "t04.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "Excellent experience from start to finish. Highly recommend this to anyone looki" + }, + { + "file": "t05.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "Absolutely disappointed with this purchase. The item arrived damaged and custome" + }, + { + "file": "t06.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The software update was installed successfully. System is functioning normally w" + }, + { + "file": "t07.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "The new interface is intuitive and the new features are genuinely useful. Great " + }, + { + "file": "t08.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "I'm very frustrated with this app. It keeps crashing and the latest update remov" + }, + { + "file": "t09.txt", + "expected": "neutral", + "difficulty": "medium", + "category": "sentiment", + "preview": "The package contains 250 units as ordered. Shipping was completed within the agr" + }, + { + "file": "t10.txt", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "preview": "Despite some initial confusion, the support team was patient and helped resolve " + }, + { + "file": "t11.txt", + "expected": "negative", + "difficulty": "medium", + "category": "sentiment", + "preview": "The documentation is outdated and incomplete. Half the examples don't work and k" + }, + { + "file": "t12.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Monthly recurring revenue increased 2.3% quarter-over-quarter. Customer churn ra" + }, + { + "file": "t13.txt", + "expected": "positive", + "difficulty": "hard", + "category": "sentiment", + "preview": "The subtle improvements to the notification system really make a difference in d" + }, + { + "file": "t14.txt", + "expected": "negative", + "difficulty": "hard", + "category": "sentiment", + "preview": "The constant notifications are disruptive and I find the new design choices ques" + }, + { + "file": "t15.txt", + "expected": "positive", + "difficulty": "easy", + "category": "sentiment", + "preview": "I'm thrilled with the results. The quality exceeded expectations and delivery wa" + }, + { + "file": "t16.txt", + "expected": "negative", + "difficulty": "easy", + "category": "sentiment", + "preview": "Worst experience ever. The product broke within a week and I couldn't get a refu" + }, + { + "file": "t17.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Temperature reading: 72 degrees Fahrenheit. Humidity: 45%. No anomalies detected" + }, + { + "file": "t18.txt", + "expected": "positive", + "difficulty": "medium", + "category": "sentiment", + "preview": "The training workshop was well-organized and the instructors were knowledgeable." + }, + { + "file": "t19.txt", + "expected": "negative", + "difficulty": "medium", + "category": "sentiment", + "preview": "The conference was overcrowded and poorly organized. Sessions started late repea" + }, + { + "file": "t20.txt", + "expected": "neutral", + "difficulty": "hard", + "category": "sentiment", + "preview": "Server uptime this month: 99.87%. Average response time: 142ms. Number of incide" + } + ] +} \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json new file mode 100644 index 000000000000..53c2d1ef300c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/schema.json @@ -0,0 +1,12 @@ +{ + "properties": { + "sentiment": { + "type": "string", + "enum": [ + "positive", + "negative", + "neutral" + ] + } + } +} \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt new file mode 100644 index 000000000000..85df534e13c1 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t01.txt @@ -0,0 +1 @@ +This product is absolutely amazing! I love it and would recommend it to everyone. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt new file mode 100644 index 000000000000..665b2b23bc14 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t02.txt @@ -0,0 +1 @@ +This service was terrible. The staff was rude, the food was cold, and I had to wait an hour. Never coming back. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt new file mode 100644 index 000000000000..465ccc8fa2e1 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t03.txt @@ -0,0 +1 @@ +The meeting is scheduled for 3 PM in conference room B. Please bring the quarterly report. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt new file mode 100644 index 000000000000..4fe834ea1977 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t04.txt @@ -0,0 +1 @@ +Excellent experience from start to finish. Highly recommend this to anyone looking to improve their workflow. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt new file mode 100644 index 000000000000..d4ad6be49851 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t05.txt @@ -0,0 +1 @@ +Absolutely disappointed with this purchase. The item arrived damaged and customer service was unresponsive. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt new file mode 100644 index 000000000000..99593ec45f93 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t06.txt @@ -0,0 +1 @@ +The software update was installed successfully. System is functioning normally with no errors reported. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt new file mode 100644 index 000000000000..ee8c7d259e6c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t07.txt @@ -0,0 +1 @@ +The new interface is intuitive and the new features are genuinely useful. Great job on the redesign! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt new file mode 100644 index 000000000000..99c217ac0c9f --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t08.txt @@ -0,0 +1 @@ +I'm very frustrated with this app. It keeps crashing and the latest update removed features I relied on. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt new file mode 100644 index 000000000000..092edc95a4ab --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t09.txt @@ -0,0 +1 @@ +The package contains 250 units as ordered. Shipping was completed within the agreed timeframe of 3-5 business days. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt new file mode 100644 index 000000000000..2b933f269759 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t10.txt @@ -0,0 +1 @@ +Despite some initial confusion, the support team was patient and helped resolve the issue quickly. Thank you! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt new file mode 100644 index 000000000000..776e2296739e --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t11.txt @@ -0,0 +1 @@ +The documentation is outdated and incomplete. Half the examples don't work and key features are undocumented. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt new file mode 100644 index 000000000000..c06f4595734d --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t12.txt @@ -0,0 +1 @@ +Monthly recurring revenue increased 2.3% quarter-over-quarter. Customer churn rate is 1.2% below industry average. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt new file mode 100644 index 000000000000..9133644f757a --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t13.txt @@ -0,0 +1 @@ +The subtle improvements to the notification system really make a difference in daily productivity. Worth the upgrade. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt new file mode 100644 index 000000000000..c08983c91a1c --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t14.txt @@ -0,0 +1 @@ +The constant notifications are disruptive and I find the new design choices questionable at best. Regretting this update. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt new file mode 100644 index 000000000000..52354cff20bc --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t15.txt @@ -0,0 +1 @@ +I'm thrilled with the results. The quality exceeded expectations and delivery was faster than promised. Five stars! \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt new file mode 100644 index 000000000000..2ed09b28b27e --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t16.txt @@ -0,0 +1 @@ +Worst experience ever. The product broke within a week and I couldn't get a refund. Stay away from this seller. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt new file mode 100644 index 000000000000..f271e1e6dbc6 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t17.txt @@ -0,0 +1 @@ +Temperature reading: 72 degrees Fahrenheit. Humidity: 45%. No anomalies detected in the past 24-hour monitoring period. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt new file mode 100644 index 000000000000..09d654200723 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t18.txt @@ -0,0 +1 @@ +The training workshop was well-organized and the instructors were knowledgeable. I learned several valuable techniques. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt new file mode 100644 index 000000000000..c519d34837db --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t19.txt @@ -0,0 +1 @@ +The conference was overcrowded and poorly organized. Sessions started late repeatedly and the venue was subpar. \ No newline at end of file diff --git a/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt new file mode 100644 index 000000000000..b787a856a306 --- /dev/null +++ b/tools/parallel-decision/examples/vision-decision-harness/tests/text_samples/t20.txt @@ -0,0 +1 @@ +Server uptime this month: 99.87%. Average response time: 142ms. Number of incidents reported: 3. All resolved within SLA. \ No newline at end of file diff --git a/tools/server/server-context.cpp b/tools/server/server-context.cpp index 41c6da3e1bbc..28a6b3b75050 100644 --- a/tools/server/server-context.cpp +++ b/tools/server/server-context.cpp @@ -9,6 +9,7 @@ #include "build-info.h" #include "common.h" +#include "base64.hpp" #include "fit.h" #include "llama.h" #include "log.h" @@ -39,6 +40,10 @@ constexpr int HTTP_POLLING_SECONDS = 1; +// Toggle debug output with LLAMA_DECISION_DEBUG env var +namespace { bool decision_debug_enabled() { static bool v = std::getenv("LLAMA_DECISION_DEBUG") != nullptr; return v; } } +#define DECISION_DEBUG(fmt, ...) do { if (decision_debug_enabled()) { fprintf(stderr, "[decision-debug] " fmt "\n", ##__VA_ARGS__); } } while(0) + static common_speculative_output_limits server_output_limits(const common_params & params) { if (params.embedding || (params.pooling_type != LLAMA_POOLING_TYPE_UNSPECIFIED && params.pooling_type != LLAMA_POOLING_TYPE_NONE)) { @@ -2387,21 +2392,99 @@ struct server_context_impl { if (!body.contains("contexts") || !body.at("contexts").is_array() || body.at("contexts").empty() || body.at("contexts").size() > 256) { throw std::invalid_argument("\"contexts\" must be an array of 1-256 strings"); } + // Optional: images array (base64-encoded), one per context, in the same order as contexts. + // An empty/null entry in images means "no image for this context". + // Each entry can also be an array of base64 strings for multi-image contexts. std::vector contexts; - for (const auto & c : body.at("contexts")) { + std::vector context_bitmaps; + bool has_images = body.contains("images") && body.at("images").is_array(); + DECISION_DEBUG("handle_decision: has_images=%d", (int)has_images); + if (has_images && body.at("images").size() != body.at("contexts").size()) { + throw std::invalid_argument("\"images\" array length must match \"contexts\" length"); + } + contexts.reserve(body.at("contexts").size()); + context_bitmaps.reserve(body.at("contexts").size()); + for (size_t i = 0; i < body.at("contexts").size(); ++i) { + const auto & c = body.at("contexts")[i]; if (!c.is_string() || c.get().empty()) { throw std::invalid_argument("every entry of \"contexts\" must be a non-empty string"); } - contexts.push_back(c.get()); + std::string ctx_text = c.get(); + mtmd::bitmaps ctx_bitmaps; + if (has_images && i < body.at("images").size()) { + // images[i] can be a single base64 string OR an array of base64 strings + const auto & img_entry = body.at("images")[i]; + std::vector img_list; + if (img_entry.is_array()) { + for (const auto & img : img_entry) { + if (img.is_string() && !img.get().empty()) { + img_list.push_back(img.get()); + } + } + } else if (img_entry.is_string() && !img_entry.get().empty()) { + img_list.push_back(img_entry.get()); + } + for (const std::string & b64 : img_list) { + DECISION_DEBUG("handle_decision: decoding image for context %zu", i); + // decode base64 image + std::string raw_b64 = b64; + // strip optional data URL prefix: "data:image/...;base64,...." + if (raw_b64.find("data:") == 0) { + auto pos = raw_b64.find(",base64,"); + if (pos != std::string::npos) { + raw_b64 = raw_b64.substr(pos + 8); + } else { + auto pos2 = raw_b64.find(","); + if (pos2 != std::string::npos) { + raw_b64 = raw_b64.substr(pos2 + 1); + } + } + } + std::string raw = base64::decode(raw_b64); + DECISION_DEBUG("handle_decision: image raw size=%zu", raw.size()); + if (!raw.empty()) { + auto out = mtmd_helper_bitmap_init_from_buf(mctx, reinterpret_cast(raw.data()), raw.size(), false, init_opt); + DECISION_DEBUG("handle_decision: bitmap init result.bitmap=%p", (void*)out.bitmap); + if (out.bitmap) { + ctx_bitmaps.entries.emplace_back(out.bitmap); + } else { + throw std::runtime_error("failed to decode image at context " + std::to_string(i)); + } + } + } + } + if (!ctx_bitmaps.entries.empty()) { + // Media markers should already be in the context text at this point. + // If the context text doesn't contain any media markers, prepend them + // (backward compatibility with single-image contexts). + // The harness can now include media markers inline in context text + // for multi-image contexts where image order matters relative to text. + const char * marker = mctx ? mtmd_get_marker(mctx) : nullptr; + if (marker && ctx_text.find(marker) == std::string::npos) { + // No media markers in text, so prepend them (legacy behavior) + std::string markers; + for (size_t j = 0; j < ctx_bitmaps.entries.size(); ++j) { + markers += marker; + } + ctx_text = markers + ctx_text; + DECISION_DEBUG("handle_decision: prepended %zu media markers", ctx_bitmaps.entries.size()); + } + // If markers are already in the text, mtmd_tokenize will find and use them + } + contexts.push_back(ctx_text); + context_bitmaps.emplace_back(std::move(ctx_bitmaps)); } if (!body.contains("schema")) { throw std::invalid_argument("\"schema\" must be provided"); } if (!decision_engine) { + DECISION_DEBUG("handle_decision: creating decision engine"); decision_engine = std::make_unique(ctx_tgt, (llama_seq_id) params_base.n_parallel, - params_base.n_seq_decision); + params_base.n_seq_decision, mctx); } + DECISION_DEBUG("handle_decision: compiling schema"); const auto cs = llama_decision::compile_schema(body.at("schema"), body.value("instructions", std::string())); + DECISION_DEBUG("handle_decision: rendering prompts"); std::string shared; std::vector dynamic; for (const auto & c : contexts) { @@ -2418,6 +2501,8 @@ struct server_context_impl { opt.tree_max = (size_t) body.value("tree_max", 128); opt.allow_cache = body.value("cache_prompt", true); + // If any context has images, pass the bitmaps through options for multimodal tokenization + opt.context_bitmaps = std::move(context_bitmaps); const auto b = decision_engine->decide_batch(shared, dynamic, cs.inputs, opt); size_t context_tokens = 0;