Repository navigation
fix(core): forward task images into the first LLM turn (multimodal passthrough) - #828
Open
louisss1016 wants to merge 1 commit into
Open
louisss1016 wants to merge 1 commit into
louisss1016 wants to merge 1 commit into
Conversation
…A-1) put_task() has always accepted images and every frontend passes them (fsapp downloads Feishu images to disk, desktop_bridge/tui_v3 forward paths), but run() never handed them to agent_runner_loop — the first user turn stayed plain text and attachments were silently dropped. - agentmain: new _multimodal_initial_content() builds Claude-style content blocks (text + base64 image blocks); run() passes them via agent_runner_loop(initial_user_content=...). Gated on NativeToolClient so other clients keep plain-text behaviour. Non-vision formats (e.g. SVG) degrade to a text path reference, and unreadable files surface as an error text block instead of vanishing. - llmcore: _to_responses_input() now converts Claude-style image blocks to input_image. Previously only OpenAI-style image_url blocks survived, so attachments died on the Responses API path even when they reached the converter. - llmcore: extract _drop_blank_text_blocks(). The old inline filter dropped every block without non-blank text — image blocks included — so NativeToolClient.chat() stripped attachments before the backend ever saw them. - desktop_bridge: remove the _patch_chat_for_images monkey-patch. It existed to work around the missing core plumbing and would now double-inject every image; attachments flow through put_task instead. Its guard tests are repointed at the core mechanism. Tests: frontends/tests/test_multimodal_image_passthrough.py (21 cases: responses converter image/url/malformed/assistant branches, blank-text filter keeping image/tool_result blocks, NativeToolClient.chat end-to-end passthrough, helper extraction via ast, run() wiring). Full suite: 294 passed with the same 3 pre-existing environmental failures as the untouched baseline (272 passed + symlink x2, release packaging). Refs lsdefine#813 (item A-1) Co-Authored-By: GiftedScout <GiftedScout@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #813 (item A-1 — multimodal images silently dropped). Follow-up to #822 (A-4/A-5).
Problem
put_task()has always accepted images and every frontend passes them (fsapp downloads Feishu images to disk and forwards the paths; desktop_bridge and tui_v3 forward paths too), butagentmain.run()never handed them toagent_runner_loop— the first user turn stayed plain text and attachments were silently dropped. The desktop bridge papers over this with abackend.askmonkey-patch; fsapp and tui have no such workaround.Even when image blocks do reach the LLM layer, two more drop points exist:
_to_responses_input()only understood OpenAI-styleimage_urlblocks — Claude-style{"type": "image", "source": {...}}blocks vanished on the Responses API path.NativeToolClient.chat()'s whitespace filter ([c for c in combined_content if c.get("text", "").strip()]) dropped every block without non-blank text, image blocks included.Fix
_multimodal_initial_content()builds the first turn's content blocks (text + base64 image blocks);run()passes them viaagent_runner_loop(initial_user_content=...). Gated onNativeToolClient— only those backends understand image blocks (native Claude API takes them as-is; the OAI paths convert them), so other clients keep plain-text behaviour. Non-vision formats (e.g. SVG) degrade to a text path reference, and unreadable files surface as an error text block instead of vanishing silently._to_responses_input()converts Claude-style image blocks (base64 and url sources) toinput_image, mirroring_msgs_claude2oai()._drop_blank_text_blocks(); only blank text blocks are dropped now, image/tool_result blocks survive._patch_chat_for_imagesmonkey-patch — with the core plumbing fixed it would double-inject every image. Attachments flow throughput_task()instead. The guard tests intest_bridge_submit.pyare repointed at the core mechanism (test_core_forwards_images_to_agent_loop,test_run_agent_turn_does_not_monkeypatch_images).Tests
New
frontends/tests/test_multimodal_image_passthrough.py(21 cases, red before the fix):_to_responses_input: base64/url image blocks ->input_image, OpenAI-styleimage_urlregression guard, assistant role and malformed blocks ignored_drop_blank_text_blocks: blank text dropped, image/tool_result blocks keptNativeToolClient.chatend-to-end: image blocks reachbackend.ask, blank text still dropped_multimodal_initial_content(extracted via ast, no import): no images / non-native client -> None, text+image blocks with verified base64 bytes, unknown extension -> image/png, SVG -> text reference, unreadable path -> error text, dict-shaped entriesrun()wiring:agent_runner_loopreceivesinitial_user_contentfed fromtask.get("images")Full suite: 294 passed, with the same 3 pre-existing environmental failures as the untouched baseline (272 passed + symlink x2 + release packaging) — zero regressions, +22 new passing tests.