You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When a tool function executed during Agent.run() dynamically loaded new tools — for
example discovering additional MCP capabilities — those tools were never visible to the
model for the rest of that run. Agent._prepare_run_context snapshotted the agent's
tools, mcp_tools, and any run-level tools= list exactly once at run start, and the
function-calling loop kept reusing that same snapshot on every subsequent model call.
Callers had no way to inject a tool mid-run: they had to discard the current response and
re-invoke run() with an updated tool list, costing an extra LLM round trip per dynamic
tool load.
Description & Review Guide
What are the major changes?
Agent._prepare_run_context now records which tool-list items (agent tools, MCP
servers, run-level tools= list) are present when the run starts, and passes the
function-invocation loop a private refresh_run_tools callback.
The loop (FunctionInvocationLayer, both the streaming and non-streaming paths) calls
that callback before every model request, merging any newly appended tools into the
run-local tools list in place. Because FunctionInvocationContext.tools is that same
list, tools added this way are visible to add_tools/remove_tools and to the tool
map on the very next turn.
A newly appended MCPTool is connected and its functions exposed the same way MCP
servers configured at run start already are (the connect-and-expose logic was
extracted into a shared helper, _append_run_tools, reused by both paths).
A duplicate tool name from the refresh is logged and skipped rather than aborting the
run, since the model may already have in-flight calls against the existing tools.
No public API changes. FunctionInvocationContext.add_tools() / remove_tools()
(Python: progressive tool exposure via FunctionInvocationContext #6233) are unaffected; this closes the gap for the other two sources the issue
described — the agent's own tool list and a run's tools= list — instead of adding a
new mechanism.
Updated the "Runtime tool changes" scenario in docs/specs/004-python-function-calling-loop.md per this repo's function-calling-loop
change process.
What is the impact of these changes?
When nothing is appended mid-run, behavior is unchanged: the loop reuses the exact same
tools list object across iterations, with no extra copying or re-resolution work beyond
one cheap comparison per iteration.
What do you want reviewers to focus on?
Whether re-resolving tool sources on every iteration (rather than only when a mutation
is detected some other way) is an acceptable cost for the common no-op case.
The duplicate-name-is-skipped-with-a-warning behavior, versus failing the run outright.
This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
This is not a breaking change. If it is a breaking change, add the breaking change label (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.
New items are refreshed in MCP-first order, but run-start resolution at lines 1639-1650 orders agent tools, then run-level tools, then MCP tools. If a tool call appends an MCP server and a local/run tool with the same name in one turn, the MCP function therefore wins this run even though the normal source precedence puts the local tool first. Preserve the run-start source order here so duplicate-name handling is deterministic across initial and refreshed tools.
The common no-op path rebuilds a list containing every source item and scans it on every model turn, so its cost is O(total tools) plus one allocation—not the “one cheap comparison” and “no extra copying” stated in the PR description. This matters for the hundreds-of-tools scenario motivating dynamic exposure; track each source list's identity/processed length and inspect only appended suffixes, or revise and justify the stated performance impact.
Clarified docstring to specify discovery order of tool sources and behavior for duplicate names.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
documentationUsage: [Issues, PRs], Target: documentation in the code base and learn docspythonUsage: [Issues, PRs], Target: Python
2 participants
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation & Context
When a tool function executed during
Agent.run()dynamically loaded new tools — forexample discovering additional MCP capabilities — those tools were never visible to the
model for the rest of that run.
Agent._prepare_run_contextsnapshotted the agent'stools,
mcp_tools, and any run-leveltools=list exactly once at run start, and thefunction-calling loop kept reusing that same snapshot on every subsequent model call.
Callers had no way to inject a tool mid-run: they had to discard the current response and
re-invoke
run()with an updated tool list, costing an extra LLM round trip per dynamictool load.
Description & Review Guide
What are the major changes?
Agent._prepare_run_contextnow records which tool-list items (agent tools, MCPservers, run-level
tools=list) are present when the run starts, and passes thefunction-invocation loop a private
refresh_run_toolscallback.FunctionInvocationLayer, both the streaming and non-streaming paths) callsthat callback before every model request, merging any newly appended tools into the
run-local tools list in place. Because
FunctionInvocationContext.toolsis that samelist, tools added this way are visible to
add_tools/remove_toolsand to the toolmap on the very next turn.
MCPToolis connected and its functions exposed the same way MCPservers configured at run start already are (the connect-and-expose logic was
extracted into a shared helper,
_append_run_tools, reused by both paths).run, since the model may already have in-flight calls against the existing tools.
FunctionInvocationContext.add_tools()/remove_tools()(Python: progressive tool exposure via FunctionInvocationContext #6233) are unaffected; this closes the gap for the other two sources the issue
described — the agent's own tool list and a run's
tools=list — instead of adding anew mechanism.
docs/specs/004-python-function-calling-loop.mdper this repo's function-calling-loopchange process.
What is the impact of these changes?
When nothing is appended mid-run, behavior is unchanged: the loop reuses the exact same
tools list object across iterations, with no extra copying or re-resolution work beyond
one cheap comparison per iteration.
What do you want reviewers to focus on?
is detected some other way) is an acceptable cost for the common no-op case.
Related Issue
Fixes #4024
Contribution Checklist
breaking changelabel (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.