Skip to content

feat: add OpenAI Responses API endpoint (/v1/responses) to m serve - #1708

Draft
markstur wants to merge 22 commits into
generative-computing:mainfrom
markstur:responses
Draft

markstur wants to merge 22 commits into
generative-computing:mainfrom
markstur:responses

Conversation

@markstur

@markstur markstur commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request

Issue

Fixes #1635

Description

Adds responses endpoint consistent with the OpenAI API and our chat/completions endpoint.
This brings the updated way of interacting which can include server-side storing of context and server-side tools.
The server-side storage is very basic in-memory with time-to-live setting. Storage that lasts across restarts and is more management options could be added later.

Testing

  • Tests added to the respective file if code was changed
  • New code has 100% coverage if code was added
  • Ensure existing tests and github automation passes (a maintainer will kick off the github automation when the rest of the PR is populated)

Attribution

  • AI coding assistants used

Adding a new component, requirement, sampling strategy, or tool?

If your PR adds or modifies one of the types below, check the matching box. A checklist of type-specific review items will be posted as a comment.

  • Component
  • Requirement
  • Sampling Strategy
  • Tool

NOTE: Please ensure you have an issue that has been acknowledged by a core contributor and routed you to open a pull request against this repository. Otherwise, please open an issue before continuing with this pull request.

Add support for the OpenAI Responses API alongside the existing
/v1/chat/completions endpoint. Both routes are registered on the same
server and share the same user-supplied serve() function.

New endpoint behaviour:
- Accepts ResponseRequest (input, instructions, tools, conversation,
  previous_response_id, max_output_tokens, temperature, etc.)
- Converts string or message-array input to the internal ChatMessage
  format; maps the `developer` role to `system`
- Returns a typed Response object with an output[] array, output_text
  convenience field, and ResponseUsage (including reasoning_tokens)
- Streaming emits semantic SSE events: response.created,
  response.in_progress, response.output_text.delta,
  response.output_text.done, response.function_call_arguments.done,
  response.completed (or response.failed on error)
- Tool calls produce ResponseFunctionCall output items
- background=true returns 400 (not yet supported)

New models in cli/serve/models.py:
  ResponseRequest, ResponseInputItem, InputContent, ResponseTool,
  WebSearchTool, FileSearchTool, MCPTool, ContextManagementConfig,
  PromptCacheOptions, Response, ResponseError

New helpers in mellea/helpers/openai_compatible_helpers.py:
  ResponseUsage, ResponseOutputItem, ResponseOutputMessage,
  ResponseFunctionCall, OutputTextContent, build_response_usage(),
  build_response_output_items()

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Update m-serve.md to mention both endpoints in the intro, add
/v1/responses to the exposed routes list, and add a "Responses API"
section with curl, OpenAI SDK, and streaming examples.

Fix the API Endpoints section in docs/examples/m_serve/README.md,
which incorrectly listed POST /generate (not a real route). Replace
it with the actual registered routes: /v1/chat/completions and
/v1/responses.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Add client_responses.py, client_responses_streaming.py, and
client_responses_tool_calling.py alongside the existing Chat
Completions clients in simple/, streaming/, and tool-calling/.
Each uses the same server program as its sibling client — no server
changes are needed to use /v1/responses.

client_responses.py: minimal smoke test using client.responses.create()
and output_text.

client_responses_streaming.py: toggleable streaming/non-streaming using
client.responses.stream() and response.output_text.delta events.

client_responses_tool_calling.py: three scenarios (weather, stock price,
tool call + follow-up) reading function_call items from the output[]
array.

Update docs/examples/m_serve/README.md to list the new files and note
that all Responses API clients share the same server program.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…s API

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Fitting the new responses paradigm, we can continue a conversation w/o passing everything back and forth.

At this time, there is no store that persists a server restart.
There is a time-to-live (TTL) option as a simplistic way to not hog memory forever.
Store, compaction, cancel, etc... would likely be future enhancements.

Assisted-by: IBM Bob
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
…letion IDs

Replace uuid.uuid4().hex[:N] with secrets.choice over ascii_letters+digits
for both response_id (resp_...) and completion_id (chatcmpl-...) generation.

The hex approach produced only lowercase a-f plus 0-9, which does not match
the mixed-case alphanumeric format used by the OpenAI API. The new IDs use
the full base-62 alphabet (0-9a-zA-Z), matching the OpenAI spec.

Also fix test_previous_response_id_echoed, which was passing an unknown ID
and expecting it to be echoed back. The correct behaviour (404 for unknown
previous_response_id) meant the endpoint returned a JSONResponse, causing
an AttributeError. The test now does a proper two-step round-trip: store a
response first, then reference it via previous_response_id.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Add expires_at: int | None to the Response model. Set to created_at +
TTL when store=True, None when store=False. Updates README to document
in-memory-only behaviour and loss on restart.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
web_search, file_search, mcp, and code_interpreter were accepted by the
schema but had no implementation, causing an internal AssertionError deep
in the backend. Now _build_model_options_from_response_request raises
ValueError for any non-function tool type, which the existing handler
converts to a 400 with an actionable message.

Remove the unused WebSearchTool, FileSearchTool, and MCPTool model
classes — the type literals on ResponseTool are sufficient for parsing
before rejection.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
stream_response_chunks ignored the request's include list and always
emitted usage in response.completed. Pass request.include through from
make_responses_endpoint and gate usage on "usage" in include.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Assisted-by: IBM Bob
Response.output was typed as Union[ResponseOutputMessage,
ResponseFunctionCall, ResponseOutputItem] but Pydantic was matching
items against the base class first, stripping the content and role
fields from serialized message items.

Reorder the Union so ResponseOutputMessage and ResponseFunctionCall
are listed before the base ResponseOutputItem, ensuring subclass
fields are preserved on serialization.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
stream_response_chunks() had two gaps:

1. response.completed emitted only {id, status, usage}; SDK clients
   expect the full Response object. Now builds a complete Response and
   serializes it as {"type": "response.completed", "response": ...}.

2. Tool call arguments were only emitted as .done with no preceding
   .delta. Now emits response.function_call_arguments.delta (full
   arguments as one chunk) before .done, which is spec-conformant when
   the backend doesn't stream arguments incrementally.

Also adds the missing SDK envelope fields to every event:
- type field matching the event name
- sequence_number counter
- response.output_item.added before content events
- response.content_part.added before text delta/done events

Passes store, ttl, and previous_response_id into the generator so
the completed envelope has the correct expires_at and
previous_response_id fields.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Supports structured output same as chat/completions.
Follows OpenAI API spec.
Does not support json_object mode yet (same as completions).

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Consistent with chat/completions implementation.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
@markstur
markstur requested a review from a team as a code owner October 5, 2026 15:59
@markstur
markstur marked this pull request as draft October 5, 2026 15:59
@markstur

markstur commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

marking as draft because some fixes are coming

Fix streaming events for OpenAI SDK spec.
Fix tool call index.

Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
@markstur markstur changed the title Responses feat: add OpenAI Responses API endpoint (/v1/responses) to m serve Oct 5, 2026
@github-actions github-actions Bot added the enhancement New feature or request label Oct 5, 2026
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>
Signed-off-by: Mark Sturdevant <mark.sturdevant@ibm.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: add OpenAI Responses API endpoint (/v1/responses) to m serve

1 participant