Skip to content

Python: [Feature] Allow tools to declare standing guidance appended to results - #8784

Open
pratik wayase (PratikWayase) wants to merge 3 commits into
microsoft:mainfrom
PratikWayase:feat/middleware-append-guidance-8757
Open

pratik wayase (PratikWayase) wants to merge 3 commits into
microsoft:mainfrom
PratikWayase:feat/middleware-append-guidance-8757

Conversation

@PratikWayase

Copy link
Copy Markdown
Contributor

Motivation & Context

This change resolves a gap opened in v1.19 regarding how tools communicate the meaning of their hidden or untrusted results to the model.

In v1.19, LabelTrackingFunctionMiddleware correctly closed a security hole by making per-item security labels restrict-only (unless carrying a private framework marker). This prevented a tool body from promoting runtime bytes (like web content) to trusted. However, this removed the only channel a tool had to provide explanatory context about its own hidden output. A hidden failed compile and a hidden clean compile look identical to the model.

Currently, integrators are using unsafe workarounds, such as stamping the private _INTERNAL_RESULT_MARKER directly or declaring the whole tool trusted and inverting label flow, which widens the attack surface. This PR provides a supported, structural mechanism for tools to declare standing guidance.

Description & Review Guide

  • What are the major changes?

    • Added support for standing_guidance: list[str] in a tool's additional_properties declaration.
    • Updated LabelTrackingFunctionMiddleware._process_result_with_embedded_labels flow to construct and append these guidance sentences as their own Content items.
    • The middleware itself constructs these items and stamps them with the framework's authoritative _INTERNAL_RESULT_MARKER and IntegrityLabel.TRUSTED, using the tool's resolved confidentiality label.
    • Added validation in _standing_guidance_items to gracefully ignore malformed (non-string/empty) entries with a warning, rather than raising exceptions.
  • What is the impact of these changes?

    • Tools can now safely append trusted context (e.g., "A result you cannot read is not a clean validation.") alongside their hidden untrusted output.
    • The text is fixed at declaration time and never passes through the tool body, so an attacker controlling runtime inputs cannot alter it. The framework acts as the producer, requiring no new authority grants to third parties.
    • No breaking changes; behavior remains identical for tools that do not declare standing_guidance.
  • What do you want reviewers to focus on?

    • The security label stamping in _standing_guidance_items: Confirm that constructing the Content in the middleware and applying _INTERNAL_RESULT_MARKER correctly aligns with the framework's authority model (similar to quarantined_llm).
    • The flow in _process_result_with_embedded_labels: Verify that injecting the guidance items into original_items before processing correctly preserves them visibly while hiding the actual untrusted tool output.

Related Issue

Fixes #8757

Contribution Checklist

  • The code builds clean without any errors or warnings
  • All unit tests pass, and I have added new tests where possible
  • The PR follows the Contribution Guidelines
  • This PR is linked to an issue and there is no other open PR for this issue (see Related Issue above).
  • This is not a breaking change. If it is a breaking change, add the breaking change label (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Trusted guidance remains mutable during execution and mishandles USER_IDENTITY principal metadata.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)
What changed in this PR

Adds trusted, tool-declared standing guidance to explain hidden security-labeled results.

Changes:

  • Appends validated guidance as authoritative trusted content.
  • Handles absent, malformed, and None results.
  • Adds focused middleware tests.
File Description
python/​packages/​core/​agent_framework/​security.py Implements guidance creation and result processing.
python/​packages/​core/​tests/​test_security.py Tests guidance behavior and validation.

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/packages/core/agent_framework/security.py Outdated
Comment thread python/packages/core/agent_framework/security.py
Comment thread python/packages/core/agent_framework/security.py Outdated
Comment thread python/packages/core/agent_framework/security.py Outdated
in ``_clone_for_scope()``, so scoped middleware instances reuse the same
frozen snapshot.
"""
key = id(function)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pratik wayase (@PratikWayase) This cache cannot safely use id(function) as a long-lived key. Python can reuse that integer after a short-lived tool is collected, so a later unrelated tool can receive the prior tool’s cached guidance; _standing_guidance_items() then stamps that stale text with the framework-owned marker and TRUSTED integrity. Please key the cache by live object identity (for example, a weak-key cache), store the frozen tuple on the registration object, or otherwise ensure entries cannot outlive the tool object.

This branch was successfully deployed

1 active deployment
github-app-auth — f893122e Deployed Sep 28, 2026 by PratikWayase via add_label #23905
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python: [Feature]: Let a tool declare standing guidance the middleware appends to its results

4 participants