Repository navigation
Conversation
…n non-UTF-8
A tool annotated `-> bytes` advertises output schema {"type":"string","format":"binary"}, but returning non-UTF-8 binary data (e.g. PNG magic bytes) crashed with PydanticSerializationError.
Root cause: _convert_to_content fed bytes to pydantic_core.to_json, which UTF-8-decodes them (fallback=str only applies to types pydantic cannot serialize), and the generated structured-output model serialized bytes with the default ser_json_bytes='utf8'. Both raise on non-UTF-8 bytes; UTF-8-decodable bytes also came back as a raw JSON string rather than base64.
Fix:
- _convert_to_content now base64-encodes bytes before the to_json branch, matching Image/Audio and the lowlevel server, so all bytes become a consistent base64 string.
- Generated output models (_create_wrapped_model, _create_model_from_class) set ser_json_bytes='base64' so the structured-content path serializes bytes as base64 instead of crashing.
- StrictJsonSchema.bytes_schema pins the advertised schema to format: binary, so ser_json_bytes='base64' does not leak base64url into outputSchema.
Only generated output models are covered; a user-defined BaseModel/TypedDict with a bytes field is unchanged (pydantic forbids a TypeAdapter config override there).
Signed-off-by: Udaya Tejas <udayatejas2004@gmail.com>
|
This PR has been closed automatically. This repo only keeps pull requests open when they come from a maintainer, or from a contributor a maintainer has assigned to the linked issue, and you aren't currently assigned to #3554. If a maintainer assigns you to #3554, this PR reopens on its own and there's nothing more you need to do here. Assignment is a maintainer call based on capacity; comments that only ask to be assigned don't factor in. What does help is engaging on the issue itself by confirming the repro, explaining why it matters for your use case, or describing the approach you'd take. You're welcome to keep pushing commits here (just avoid force-pushing, since GitHub can't reopen a rewritten branch), but that on its own won't get the PR reviewed or the issue assigned, and realistically most auto-closed PRs stay closed. There's no need to open a new PR either way. CONTRIBUTING.md has the full reasoning, but in short:
Maintainers: reopen, remove |
Closes #3554.
Problem
A tool annotated
-> bytesadvertisesoutputSchema {"result": {"type": "string", "format": "binary"}}, but when it returns non-UTF-8 binary data (PNG magic bytes, a zip, a PDF, …) the call crashes withPydanticSerializationErrorinside_convert_to_content; the client only getsis_error: truewith a generic message and the payload is lost. UTF-8-decodable bytes didn't crash but came back as a raw string rather than base64.Root cause
_convert_to_content(func_metadata.py) fed the result topydantic_core.to_json(result, fallback=str, …), which JSON-encodesbytesby UTF-8-decoding them —fallbackonly applies to types pydantic can't serialize, so non-UTF-8bytesraise. The generated structured-output model also serialized bytes with the defaultser_json_bytes="utf8", hitting the same failure on the structured path.Fix
Base64-encode
bytesfor JSON, matching the convention already used byImage/Audio(utilities/types.py) and the lowlevel server (server.py,base64.b64encode(...).decode()):_convert_to_contentbase64-encodesbytesbefore theto_jsonbranch, so all bytes become a consistent base64 string instead of crashing or leaking a raw decode._create_wrapped_model,_create_model_from_class) setser_json_bytes="base64"so the structured-content path serializes bytes as base64.StrictJsonSchema.bytes_schemapins the advertised schema to{"type": "string", "format": "binary"}, so enablingser_json_bytes="base64"does not leakformat: base64urlintooutputSchema— the advertised schema is unchanged.Scope note: this covers the direct
-> bytesreturn (the issue's primary case) and ordinary classes converted via_create_model_from_class. A user-definedBaseModel/TypedDictwith abytesfield is intentionally not changed here: pydantic forbids aTypeAdapter(config=...)override for those types (PydanticUserError: Cannot use config when the type is a BaseModel, dataclass or TypedDict), so fixing it would require mutating/rebuilding the user's model — a larger, separate change. Happy to follow up if maintainers want that path addressed too.Tests
Added 4 regression tests in
tests/server/mcpserver/test_func_metadata.py: non-UTF-8 bytes base64-encode in_convert_to_content; UTF-8-decodable bytes also return base64 (consistent); a-> bytestool returns base64 in both unstructured and structured output with theformat: binaryschema asserted; a bytes field in an ordinary-class output model serializes to base64 and decodes back. The 4 tests fail onmainand pass with the fix. Fulltests/server/mcpserver/suite: 646 passed; ruff + pyright clean.