Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/specs/feature-usage-bit-registry.md
Original file line number Diff line number Diff line change
Expand Up @@ -158,7 +158,7 @@ only to approved first-party endpoints.
| 56 | `openai` | OpenAI clients | `agent_framework_openai` |
| 57 | `anthropic` | Anthropic clients | `agent_framework_anthropic` |
| 58 | `bedrock` | AWS Bedrock clients | `agent_framework_bedrock` |
| 59 | `gemini` | Gemini chat client | `agent_framework_gemini` |
| 59 | `gemini` | Gemini chat and embedding clients | `agent_framework_gemini` |
| 60 | `mistral` | Mistral embedding client | `agent_framework_mistral` |
| 61 | `ollama` | Ollama clients | `agent_framework_ollama` |
| 62 | `claude` | Claude Agent SDK agent | `agent_framework_claude` |
Expand Down
4 changes: 2 additions & 2 deletions python/.env.example
Original file line number Diff line number Diff line change
Expand Up @@ -41,8 +41,8 @@ COPILOTSTUDIOAGENT__AGENTAPPID=""
ANTHROPIC_API_KEY=""
ANTHROPIC_MODEL=""
# Google Gemini
GEMINI_API_KEY=""
GEMINI_MODEL=""
GOOGLE_API_KEY=""
GOOGLE_MODEL=""
# Ollama
OLLAMA_ENDPOINT=""
OLLAMA_MODEL=""
Expand Down
8 changes: 8 additions & 0 deletions python/packages/core/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,14 @@ The vector store API is experimental under the shared `VECTOR_STORES` feature ID
deserializes results without interpreting thresholds or re-filtering returned scores. Connectors own scoring,
filter execution, score thresholds (including provider-defined/default metrics), and paging. Use native backend
execution where available, otherwise an explicit connector-local fallback or reject unsupported options
- **Embedding request options** - `upsert` accepts either flat `embeddings_options` for
all generated vector fields or `embeddings_options_by_field` keyed by logical field name,
never both. `search` and `create_vector_search_tool` accept flat `embeddings_options`
for local query generation. Core supplies declared field dimensions and rejects
conflicting values before embedding; search ignores these options when a
precomputed vector is supplied. `create_upsert_tool` and
`VectorCollectionContextProvider` forward the same operation-specific settings
to their generated tools.
- **`create_vector_search_tool`** - Creates an agent tool from any `SupportsVectorSearch` implementation
- **`create_upsert_tool` / `create_get_tool` / `create_delete_tool`** - Create agent tools for collection CRUD;
upsert and delete require approval by default, while get does not. Auto-generated keys are omitted from upsert
Expand Down
193 changes: 174 additions & 19 deletions python/packages/core/agent_framework/_vectors.py

Large diffs are not rendered by default.

4 changes: 3 additions & 1 deletion python/packages/core/agent_framework/gemini/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,11 @@
_IMPORTS: dict[str, tuple[str, str]] = {
"GeminiChatClient": ("agent_framework_gemini", "agent-framework-gemini"),
"GeminiChatOptions": ("agent_framework_gemini", "agent-framework-gemini"),
"GeminiSettings": ("agent_framework_gemini", "agent-framework-gemini"),
"GeminiEmbeddingClient": ("agent_framework_gemini", "agent-framework-gemini"),
Comment thread
eavanvalkenburg marked this conversation as resolved.
"GeminiEmbeddingOptions": ("agent_framework_gemini", "agent-framework-gemini"),
"GoogleGeminiSettings": ("agent_framework_gemini", "agent-framework-gemini"),
"RawGeminiChatClient": ("agent_framework_gemini", "agent-framework-gemini"),
"RawGeminiEmbeddingClient": ("agent_framework_gemini", "agent-framework-gemini"),
"ThinkingConfig": ("agent_framework_gemini", "agent-framework-gemini"),
}

Expand Down
8 changes: 6 additions & 2 deletions python/packages/core/agent_framework/gemini/__init__.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -3,17 +3,21 @@
from agent_framework_gemini import (
GeminiChatClient,
GeminiChatOptions,
GeminiSettings,
GeminiEmbeddingClient,
GeminiEmbeddingOptions,
GoogleGeminiSettings,
RawGeminiChatClient,
RawGeminiEmbeddingClient,
ThinkingConfig,
)

__all__ = [
"GeminiChatClient",
"GeminiChatOptions",
"GeminiSettings",
"GeminiEmbeddingClient",
"GeminiEmbeddingOptions",
"GoogleGeminiSettings",
"RawGeminiChatClient",
"RawGeminiEmbeddingClient",
"ThinkingConfig",
]
6 changes: 5 additions & 1 deletion python/packages/core/tests/core/test_gemini_namespace.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,12 +13,16 @@ def test_gemini_namespace_dir_lists_lazy_exports() -> None:
for expected in (
"GeminiChatClient",
"GeminiChatOptions",
"GeminiSettings",
"GeminiEmbeddingClient",
"GeminiEmbeddingOptions",
"GoogleGeminiSettings",
"RawGeminiChatClient",
"RawGeminiEmbeddingClient",
"ThinkingConfig",
):
assert expected in names
assert "GeminiSettings" not in names
assert "GeminiEmbeddingSettings" not in names


def test_gemini_namespace_lazy_loads_known_attribute(monkeypatch: pytest.MonkeyPatch) -> None:
Expand Down
206 changes: 204 additions & 2 deletions python/packages/core/tests/core/test_vectors.py
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ class MockEmbeddingClient(BaseEmbeddingClient):
def __init__(self) -> None:
super().__init__()
self.values: list[Any] = []
self.options: EmbeddingGenerationOptions | None = None
self.options: dict[str, Any] | None = None

async def get_embeddings(
self,
Expand All @@ -100,7 +100,7 @@ async def get_embeddings(
options: EmbeddingGenerationOptions | None = None,
) -> GeneratedEmbeddings[list[float]]:
self.values = list(values)
self.options = options
self.options = dict(options) if options is not None else None
return GeneratedEmbeddings([Embedding(vector=[float(len(str(value))), 0.5]) for value in values])


Expand Down Expand Up @@ -737,6 +737,83 @@ async def test_collection_serializes_records_and_generates_vectors() -> None:
assert embedding_client.options == {"dimensions": 2}


async def test_upsert_embedding_options_are_merged_with_field_dimensions() -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
caller_options: dict[str, Any] = {
"task_type": "RETRIEVAL_DOCUMENT",
"extra_parameters": {"provider_flag": True},
"dimensions": 2,
}

await collection.upsert([Record("one", "first", "document")], embeddings_options=caller_options)

assert embedding_client.options == caller_options
assert embedding_client.values == ["document"]
assert embedding_client.options is not None
embedding_client.options["extra_parameters"]["provider_flag"] = False
assert caller_options["extra_parameters"] == {"provider_flag": True}
assert collection.records["one"]["vector"] == [8.0, 0.5]

embedding_client.values.clear()
with pytest.raises(ValueError, match="must match its declared dimensions"):
await collection.upsert([Record("two", "second", "document")], embeddings_options={"dimensions": 3})
assert embedding_client.values == []
assert "two" not in collection.records


@pytest.mark.parametrize("invalid_dimensions", [True, 2.0, None])
async def test_upsert_rejects_invalid_dimensions_even_if_equal(invalid_dimensions: Any) -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
with pytest.raises(ValueError, match="must match its declared dimensions"):
await collection.upsert(
[Record("one", "body", "document")], embeddings_options={"dimensions": invalid_dimensions}
)
assert embedding_client.values == []


async def test_upsert_field_options_are_isolated_and_prevalidated() -> None:
first_client = MockEmbeddingClient()
second_client = MockEmbeddingClient()
definition = VectorStoreCollectionDefinition([
VectorStoreField("key", name="id"),
VectorStoreField("vector", name="first", dimensions=2, embedding_generator=first_client),
VectorStoreField("vector", name="second", dimensions=2, embedding_generator=second_client),
])
handler = VectorStoreRecordHandler(dict, definition=definition)
record = {"id": "one", "first": "first text", "second": "second text"}

await handler.serialize(record, embeddings_options={"task_type": "RETRIEVAL_DOCUMENT"})
assert first_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}
assert second_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}

by_field = {"first": {"task_type": "RETRIEVAL_DOCUMENT"}, "second": {"task_type": "RETRIEVAL_QUERY"}}
await handler.serialize(record, embeddings_options_by_field=by_field)
assert first_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}
assert second_client.options == {"task_type": "RETRIEVAL_QUERY", "dimensions": 2}
assert by_field == {"first": {"task_type": "RETRIEVAL_DOCUMENT"}, "second": {"task_type": "RETRIEVAL_QUERY"}}

first_client.values.clear()
second_client.values.clear()
with pytest.raises(ValueError, match="second.*declared dimensions"):
await handler.serialize(record, embeddings_options_by_field={"second": {"dimensions": 3}})
assert first_client.values == second_client.values == []

with pytest.raises(ValueError, match="either embeddings_options or embeddings_options_by_field"):
await handler.serialize(record, embeddings_options={}, embeddings_options_by_field={})
with pytest.raises(ValueError, match="non-generated vector field"):
await handler.serialize(record, embeddings_options_by_field={"missing": {}})
with pytest.raises(ValueError, match="non-generated vector field"):
await handler.serialize(
record,
generate_vectors=["first"],
embeddings_options_by_field={"second": {}},
)
with pytest.raises(ValueError, match="require records and generated vector fields"):
await handler.serialize(record, generate_vectors=False, embeddings_options={})


async def test_collection_empty_upsert_skips_embedding_generation() -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
Expand Down Expand Up @@ -1053,6 +1130,36 @@ async def test_vector_search_generates_query_vector_and_forwards_threshold() ->
assert responses[1]["score"] == 0.4


async def test_vector_search_uses_provider_options_and_declared_dimensions() -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
options = {"task_type": "RETRIEVAL_QUERY"}

await collection.search("find this", embeddings_options=options)

assert embedding_client.options == {"task_type": "RETRIEVAL_QUERY", "dimensions": 2}
assert options == {"task_type": "RETRIEVAL_QUERY"}
assert collection.last_search_vector == [9.0, 0.5]

embedding_client.values.clear()
with pytest.raises(ValueError, match="must match its declared dimensions"):
await collection.search("find this", embeddings_options={"dimensions": 3})
assert embedding_client.values == []

await collection.search(vector=[1.0, 0.0], embeddings_options={"dimensions": 3, **options})
assert collection.last_search_vector == [1.0, 0.0]
assert embedding_client.values == []

await collection.search("keyword terms", vector=[0.0, 1.0], embeddings_options={"dimensions": 3})
assert collection.last_search_values == "keyword terms"
assert collection.last_search_vector == [0.0, 1.0]
assert embedding_client.values == []

await MockCollection().search(vector=[1.0, 0.0], embeddings_options=options)
with pytest.raises(ValueError, match="requires an embedding generator"):
await MockCollection().search("find this", embeddings_options=options)


async def test_keyword_hybrid_search_uses_single_search_method() -> None:
collection = MockCollection()

Expand Down Expand Up @@ -1333,6 +1440,44 @@ async def test_create_search_tool_returns_mapped_results() -> None:
assert result[0].text == "one:0.9"


@pytest.mark.parametrize("search_type", ["vector", "keyword_hybrid"])
async def test_create_search_tool_forwards_embedding_options(search_type: SearchType) -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
options = {"task_type": "RETRIEVAL_QUERY"}

tool = create_vector_search_tool(
collection,
search_type=search_type,
embeddings_options=options,
top=1,
)
options["task_type"] = "RETRIEVAL_DOCUMENT"

await tool(query="find this")

assert embedding_client.values == ["find this"]
assert embedding_client.options == {"task_type": "RETRIEVAL_QUERY", "dimensions": 2}
assert collection.last_search_values == "find this"
assert collection.last_search_vector == [9.0, 0.5]
assert collection.last_search_type == search_type
assert set(tool.parameters()["properties"]) == {"query"}


async def test_create_search_tool_rejects_unusable_embedding_options() -> None:
tool = create_vector_search_tool(MockCollection(), embeddings_options={"task_type": "RETRIEVAL_QUERY"})
with pytest.raises(ValueError, match="requires an embedding generator"):
await tool(query="find this")

embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
tool = create_vector_search_tool(collection, embeddings_options={"dimensions": 3})
with pytest.raises(ValueError, match="must match its declared dimensions"):
await tool(query="find this")
assert embedding_client.values == []
assert collection.last_search_type is None


async def test_create_search_tool_supports_declared_filter_parameters() -> None:
collection = MockCollection()
collection.records["one"] = {"record_id": "one", "body": "first", "vector": [1.0, 0.0]}
Expand Down Expand Up @@ -2595,6 +2740,33 @@ async def test_vector_crud_tools_round_trip_records() -> None:
assert await get_tool.invoke(arguments={"keys": ["one"]}, skip_parsing=True) == {"records": []}


async def test_upsert_tool_forwards_provider_embedding_options() -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
options = {"task_type": "RETRIEVAL_DOCUMENT"}
upsert_tool = create_upsert_tool(collection, embeddings_options=options)
options["task_type"] = "RETRIEVAL_QUERY"

await upsert_tool.invoke(
arguments={"records": [{"id": "one", "text": "body", "vector": "document"}]},
skip_parsing=True,
)

assert embedding_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}
assert collection.records["one"]["vector"] == [8.0, 0.5]
with pytest.raises(ValueError, match="either embeddings_options or embeddings_options_by_field"):
create_upsert_tool(collection, embeddings_options={}, embeddings_options_by_field={})

by_field = {"vector": {"task_type": "RETRIEVAL_DOCUMENT"}}
by_field_tool = create_upsert_tool(collection, embeddings_options_by_field=by_field)
by_field["vector"]["task_type"] = "RETRIEVAL_QUERY"
await by_field_tool.invoke(
arguments={"records": [{"id": "two", "text": "body", "vector": "more text"}]},
skip_parsing=True,
)
assert embedding_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}


async def test_vector_crud_tools_support_auto_generated_keys_when_model_can_omit_them() -> None:
definition = VectorStoreCollectionDefinition(
[
Expand Down Expand Up @@ -2740,6 +2912,36 @@ def test_vector_collection_context_provider_configures_tools_and_approvals() ->
assert all(tool.approval_mode == "always_require" for tool in require_all.tools)


async def test_vector_collection_context_provider_forwards_embedding_options() -> None:
embedding_client = MockEmbeddingClient()
collection = MockCollection(embedding_generator=embedding_client)
provider = VectorCollectionContextProvider(
collection,
scope_filter=None,
upsert_embeddings_options={"task_type": "RETRIEVAL_DOCUMENT"},
search_embeddings_options={"task_type": "RETRIEVAL_QUERY"},
)
tools = {tool.name: tool for tool in provider.tools}

await tools["upsert"].invoke(
arguments={"records": [{"id": "one", "text": "body", "vector": "document"}]},
skip_parsing=True,
)
assert embedding_client.options == {"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 2}

await tools["search"].invoke(arguments={"query": "find this"}, skip_parsing=True)
assert embedding_client.options == {"task_type": "RETRIEVAL_QUERY", "dimensions": 2}

with pytest.raises(ValueError, match="Upsert embedding options require include_upsert_tool=True"):
VectorCollectionContextProvider(
collection, scope_filter=None, include_upsert_tool=False, upsert_embeddings_options={}
)
with pytest.raises(ValueError, match="Search embedding options require include_search_tool=True"):
VectorCollectionContextProvider(
collection, scope_filter=None, include_search_tool=False, search_embeddings_options={}
)


async def test_vector_collection_context_provider_adds_attributed_context() -> None:
collection = MockCollection(embedding_generator=MockEmbeddingClient())
details_tool = create_vector_search_tool(collection, name="search_details")
Expand Down
28 changes: 25 additions & 3 deletions python/packages/gemini/AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,25 @@
# Gemini Package (agent-framework-gemini)

Integration with Google's Gemini Developer API and Vertex AI via the `google-genai` SDK.
Integration with Google's Gemini Developer API and Enterprise (Vertex AI) via the `google-genai` SDK.
The shared `_sdk_client.create_genai_client` resolves authentication, backend mode,
and service URL for chat and embeddings.

## Core Classes

- **`RawGeminiChatClient`** - Lightweight chat client without any layers, for custom pipeline composition
- **`GeminiChatClient`** - Full-featured chat client with function invocation, middleware, and telemetry
- **`GeminiChatOptions`** - Options TypedDict for Gemini-specific parameters
- **`GeminiSettings`** - Settings loaded from environment variables
- **`GoogleGeminiSettings`** - SDK-standard `GOOGLE_*` settings loaded from environment variables
- **`GoogleGeminiSettings`** - Shared `GOOGLE_*` environment settings for chat and embeddings
- **`ThinkingConfig`** - Configuration for extended thinking
- **`RawGeminiEmbeddingClient`** - Text and multimodal embeddings without telemetry
- **`GeminiEmbeddingClient`** - Text and multimodal embeddings with telemetry (defaults to stable `gemini-embedding-2`)
- **`GeminiEmbeddingOptions`** - Per-call embedding model, dimensions, text task, and document title

`GeminiEmbeddingClient` supports only `gemini-embedding-2` and `gemini-embedding-2-preview`,
requiring per-call task instructions for text strings. Multimodal Google SDK `Content` or
media `Part` inputs receive no task prefix, even with mixed text-and-media parts. Text-only
SDK content is rejected so callers cannot bypass the task requirement. Enterprise accepts
one content per request; the client splits batches there while keeping input order.

## Gemini-specific Options

Expand All @@ -34,3 +44,15 @@ from agent_framework.gemini import GeminiChatClient
client = GeminiChatClient(model="gemini-2.5-flash")
response = await client.get_response([Message(role="user", contents=[Content.from_text("Hello")])])
```

```python
from agent_framework.gemini import GeminiEmbeddingClient

client = GeminiEmbeddingClient()
try:
result = await client.get_embeddings(
["A document"], options={"task_type": "RETRIEVAL_DOCUMENT", "dimensions": 768}
)
finally:
await client.close()
```
Loading
Loading