Skip to content

server : accept images on /decision, add harness example - #17

Open
graydini wants to merge 3 commits into
thecodacus:parallel-decisionfrom
graydini:multimodal-decision
Open

graydini wants to merge 3 commits into
thecodacus:parallel-decisionfrom
graydini:multimodal-decision

Conversation

@graydini

@graydini graydini commented Sep 28, 2026 •

Copy link
Copy Markdown

server : accept images on /decision, add harness example

What

POST /decision gains an optional images field. Each entry maps 1:1 to a
context and carries base64 image data, or an array of base64 strings when one
context should see several images. A data URL prefix is accepted and stripped.
The decoded bitmaps reach the decision engine through
options::context_bitmaps.

Media markers are prepended only when the context text does not already contain
them. A caller that interleaves markers with its own labels keeps control of
where each image lands; a caller that just wants "this image, this question"
keeps the old behaviour unchanged.

Engine

A prompt_part is now either a list of text tokens or a media chunk.
tokenize_mm expands media markers into chunks and reports text tokens with
LLAMA_TOKEN_NULL at the image positions. decide_batch turns that into
interleaved parts, and decode_parts encodes chunks through the mtmd batch API
(mtmd_batch_init / add_chunk / encode / get_output_embd) before the
surrounding text, which is the path the completion endpoint already uses.

Models that reject a second clip chunk with "batch too large" fall back to
encoding one chunk at a time. SmolVLM2 needs that fallback; Qwen2.5-VL takes
the batch path.

LLAMA_DECISION_DEBUG in the environment traces this path.

Example

examples/vision-decision-harness/ is a Flask UI that drives the endpoint over
images and text: folder upload, batch runs with SSE progress, image selection
across a folder, a text classification suite, and snippet export. It is an
example, not a dependency, and nothing in the server links against it.

Notes

  • Vision encoding and branch scoring still complete in one batched pass per
    group. There is no autoregressive loop.
  • Text-only decisions never enter the mtmd path. The batch is only constructed
    when a context actually has a chunk.
  • Verified on Qwen2.5-VL-7B and SmolVLM2-500M. Multi-image selection over a
    12-image folder picks the intended image; the text suite runs 20 files at
    ~80% on the 7B.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: fae121c8-9f37-4ad9-954b-9a571c8527ba

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added documentation Improvements or additions to documentation server mtmd labels Sep 28, 2026
server : add images to the /decision endpoint

Add an optional "images" field to POST /decision. Each entry maps 1:1
to a context and holds base64 image data, or an array of base64 strings
when one context carries several images. A data URL prefix is accepted
and stripped. The bitmaps are passed to the decision engine through
options::context_bitmaps.

Media markers are only prepended when the context text does not already
contain them, so a caller that interleaves markers with its own labels
keeps control of where each image lands.

parallel-decision : score media chunks in the same batch

A prompt part is now either text tokens or a media chunk. tokenize_mm
expands media markers into chunks and reports text tokens with
LLAMA_TOKEN_NULL at the image positions; decide_batch turns that into
interleaved parts and decode_parts encodes chunks through the mtmd batch
API before the surrounding text.

Models that reject a second clip chunk with "batch too large" fall back
to encoding one chunk at a time, which is what SmolVLM2 needs.

Set LLAMA_DECISION_DEBUG in the environment to trace this path.

examples : add vision decision harness

A Flask UI that drives /decision over images and text: folder upload,
batch runs with SSE progress, image selection over a folder, a text
classification suite, and snippet export.
graydini added 2 commits September 29, 2026 02:22
Link parallel-decision from the root README so the feature is reachable
from the front page instead of only from its own directory.

Explain that contexts can carry images, and that chunks are encoded in the
same batched pass that scores the branches, so a decision over a
screenshot stays a single pass rather than a run of generate calls.

Add an Images section with the request shape, the positional rule for
the images array, the multi-image form, and how media markers decide
where each image lands.

Fix the endpoint path: the docs said /v1/decision, the server serves
POST /decision.
Promote the decision branch from a footnote to the main topic. The
feature was previously two bullets under Documentation, which put the
mechanism, the vision work, and the harness behind the generic tool
list.

Constrained decisions now comes first, with what the batched scoring
does, a runnable server command, a request and a response, and the
positional rule for images. The vision harness gets its own section
rather than a list item, and the debug toggle is documented next to it.

The upstream content keeps its own heading below a divider, with the
section levels demoted one step so the hierarchy stays correct.

The example response is a real one from the server.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation mtmd server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant