Skip to content

Python: refresh agent-level run tools before each function-calling turn - #8854

Open
Sivakumar Mahalingam (sivakumar-mahalingam) wants to merge 5 commits into
microsoft:mainfrom
sivakumar-mahalingam:fix/4024-dynamic-tool-injection
Open

Sivakumar Mahalingam (sivakumar-mahalingam) wants to merge 5 commits into
microsoft:mainfrom
sivakumar-mahalingam:fix/4024-dynamic-tool-injection

Conversation

@sivakumar-mahalingam

Copy link
Copy Markdown

Motivation & Context

When a tool function executed during Agent.run() dynamically loaded new tools — for
example discovering additional MCP capabilities — those tools were never visible to the
model for the rest of that run. Agent._prepare_run_context snapshotted the agent's
tools, mcp_tools, and any run-level tools= list exactly once at run start, and the
function-calling loop kept reusing that same snapshot on every subsequent model call.

Callers had no way to inject a tool mid-run: they had to discard the current response and
re-invoke run() with an updated tool list, costing an extra LLM round trip per dynamic
tool load.

Description & Review Guide

  • What are the major changes?

    • Agent._prepare_run_context now records which tool-list items (agent tools, MCP
      servers, run-level tools= list) are present when the run starts, and passes the
      function-invocation loop a private refresh_run_tools callback.
    • The loop (FunctionInvocationLayer, both the streaming and non-streaming paths) calls
      that callback before every model request, merging any newly appended tools into the
      run-local tools list in place. Because FunctionInvocationContext.tools is that same
      list, tools added this way are visible to add_tools/remove_tools and to the tool
      map on the very next turn.
    • A newly appended MCPTool is connected and its functions exposed the same way MCP
      servers configured at run start already are (the connect-and-expose logic was
      extracted into a shared helper, _append_run_tools, reused by both paths).
    • A duplicate tool name from the refresh is logged and skipped rather than aborting the
      run, since the model may already have in-flight calls against the existing tools.
    • No public API changes. FunctionInvocationContext.add_tools() / remove_tools()
      (Python: progressive tool exposure via FunctionInvocationContext #6233) are unaffected; this closes the gap for the other two sources the issue
      described — the agent's own tool list and a run's tools= list — instead of adding a
      new mechanism.
    • Updated the "Runtime tool changes" scenario in
      docs/specs/004-python-function-calling-loop.md per this repo's function-calling-loop
      change process.
  • What is the impact of these changes?
    When nothing is appended mid-run, behavior is unchanged: the loop reuses the exact same
    tools list object across iterations, with no extra copying or re-resolution work beyond
    one cheap comparison per iteration.

  • What do you want reviewers to focus on?

    • Whether re-resolving tool sources on every iteration (rather than only when a mutation
      is detected some other way) is an acceptable cost for the common no-op case.
    • The duplicate-name-is-skipped-with-a-warning behavior, versus failing the run outright.

Related Issue

Fixes #4024

Contribution Checklist

  • The code builds clean without any errors or warnings
  • All unit tests pass, and I have added new tests where possible
  • The PR follows the Contribution Guidelines
  • This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
  • This is not a breaking change. If it is a breaking change, add the breaking change label (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.

@sivakumar-mahalingam

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

A duplicate tool currently prevents valid tools from the same refresh batch from being exposed.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Refreshes agent and run-level tools before each function-calling turn.

Changes:

  • Adds streaming and non-streaming tool refresh callbacks.
  • Supports newly appended tools and MCP servers.
  • Adds regression tests and updates the function-loop specification.
File Description
python/​packages/​core/​agent_framework/​_agents.py Tracks and resolves newly added tools.
python/​packages/​core/​agent_framework/​_tools.py Refreshes tools before model requests.
python/​packages/​core/​tests/​core/​test_agents.py Tests dynamic tool behavior.
docs/​specs/​004-python-function-calling-loop.md Documents runtime tool refresh behavior.

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/packages/core/agent_framework/_tools.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Refreshed MCP tools use inconsistent source precedence when duplicate names are added simultaneously.

Review effort: Balanced
Findings: None

Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Preserve source precedence when resolving duplicate tools at run start

python/​packages/​core/​agent_framework/​_agents.py:1494

New items are refreshed in MCP-first order, but run-start resolution at lines 1639-1650 orders agent tools, then run-level tools, then MCP tools. If a tool call appends an MCP server and a local/run tool with the same name in one turn, the MCP function therefore wins this run even though the normal source precedence puts the local tool first. Preserve the run-start source order here so duplicate-name handling is deterministic across initial and refreshed tools.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The no-op refresh scans and copies all source items each turn despite the stated cost, and source-precedence documentation overstates actual behavior.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity No-op refresh still scans and copies all tools

python/​packages/​core/​agent_framework/​_agents.py:1669

The common no-op path rebuilds a list containing every source item and scans it on every model turn, so its cost is O(total tools) plus one allocation—not the “one cheap comparison” and “no extra copying” stated in the PR description. This matters for the hundreds-of-tools scenario motivating dynamic exposure; track each source list's identity/processed length and inspect only appended suffixes, or revise and justify the stated performance impact.

Comment thread python/packages/core/agent_framework/_agents.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The no-op path scans all tools each turn, and disconnected dynamic MCP lifecycle behavior lacks coverage.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
Resolved since last review (1)

Comment thread python/packages/core/agent_framework/_agents.py Outdated
Comment thread python/packages/core/tests/core/test_agents.py

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It modifies shared streaming, function-invocation, and MCP lifecycle paths that warrant final human validation.

Review effort: Balanced
Findings: None

Resolved since last review (2)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Function-loop and MCP lifecycle changes are high-risk and require final human validation.

Review effort: Balanced
Findings: None

This branch was successfully deployed

1 active deployment
github-app-auth — 6ea3c505 Deployed Sep 29, 2026 by sivakumar-mahalingam via add_label #24028
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Usage: [Issues, PRs], Target: documentation in the code base and learn docs python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Support dynamic tool injection during ChatAgent.run() execution

2 participants