Python: [Feature] Allow tools to declare standing guidance appended to results - #8784
pratik wayase (PratikWayase) wants to merge 3 commits into
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Trusted guidance remains mutable during execution and mishandles USER_IDENTITY principal metadata.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Adds trusted, tool-declared standing guidance to explain hidden security-labeled results.
Changes:
- Appends validated guidance as authoritative trusted content.
- Handles absent, malformed, and
Noneresults. - Adds focused middleware tests.
| File | Description |
|---|---|
python/packages/core/agent_framework/security.py |
Implements guidance creation and result processing. |
python/packages/core/tests/test_security.py |
Tests guidance behavior and validation. |
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
| in ``_clone_for_scope()``, so scoped middleware instances reuse the same | ||
| frozen snapshot. | ||
| """ | ||
| key = id(function) |
There was a problem hiding this comment.
pratik wayase (@PratikWayase) This cache cannot safely use id(function) as a long-lived key. Python can reuse that integer after a short-lived tool is collected, so a later unrelated tool can receive the prior tool’s cached guidance; _standing_guidance_items() then stamps that stale text with the framework-owned marker and TRUSTED integrity. Please key the cache by live object identity (for example, a weak-key cache), store the frozen tuple on the registration object, or otherwise ensure entries cannot outlive the tool object.


Motivation & Context
This change resolves a gap opened in v1.19 regarding how tools communicate the meaning of their hidden or untrusted results to the model.
In v1.19,
LabelTrackingFunctionMiddlewarecorrectly closed a security hole by making per-item security labels restrict-only (unless carrying a private framework marker). This prevented a tool body from promoting runtime bytes (like web content) totrusted. However, this removed the only channel a tool had to provide explanatory context about its own hidden output. A hidden failed compile and a hidden clean compile look identical to the model.Currently, integrators are using unsafe workarounds, such as stamping the private
_INTERNAL_RESULT_MARKERdirectly or declaring the whole tool trusted and inverting label flow, which widens the attack surface. This PR provides a supported, structural mechanism for tools to declare standing guidance.Description & Review Guide
What are the major changes?
standing_guidance: list[str]in a tool'sadditional_propertiesdeclaration.LabelTrackingFunctionMiddleware._process_result_with_embedded_labelsflow to construct and append these guidance sentences as their ownContentitems._INTERNAL_RESULT_MARKERandIntegrityLabel.TRUSTED, using the tool's resolvedconfidentialitylabel._standing_guidance_itemsto gracefully ignore malformed (non-string/empty) entries with a warning, rather than raising exceptions.What is the impact of these changes?
standing_guidance.What do you want reviewers to focus on?
_standing_guidance_items: Confirm that constructing theContentin the middleware and applying_INTERNAL_RESULT_MARKERcorrectly aligns with the framework's authority model (similar toquarantined_llm)._process_result_with_embedded_labels: Verify that injecting the guidance items intooriginal_itemsbefore processing correctly preserves them visibly while hiding the actual untrusted tool output.Related Issue
Fixes #8757
Contribution Checklist
breaking changelabel (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and title prefix in sync automatically.