Skip to content

Stabilise generated schema ordering - #272

Open
akbashev wants to merge 2 commits into
huggingface:mainfrom
akbashev:fix/prompt-cache-stability
Open

akbashev wants to merge 2 commits into
huggingface:mainfrom
akbashev:fix/prompt-cache-stability

Conversation

@akbashev

Copy link
Copy Markdown

Something I've encountered in my project:

Ollama can reuse a cached prompt prefix when the next request begins with the same serialised content. Changes in JSON key or array order can break that match, even when the schemas mean the same thing.

AnyLanguageModel stores required property names in a Set. Encoding that set directly can produce different required array orders across requests. Sorting property names and required values makes generated schemas deterministic. Sorting JSON object keys in the Ollama request encoder keeps the final request serialization stable for both streaming and non-streaming calls.

Observed in Ollama’s logs:

  • Before: 939 of 6,753 prompt tokens reused (14%); 6,749 tokens evaluated in 37.4 seconds.
  • After: 7,462 of 7,549 tokens reused (99%); 39 tokens evaluated in 0.63 seconds.

These are observations from separate runs, not a controlled benchmark. They show the intended effect: equivalent schema content now produces a much more reusable prompt prefix.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

$defs ordering remains nondeterministic, and request serialization coverage should be added.

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

This pull request stabilizes generated schema and Ollama request serialization to improve prompt-prefix cache reuse.

Changes:

  • Sorts schema properties and required fields.
  • Uses sorted JSON keys for streaming and non-streaming Ollama requests.
File Summary
Sources/​AnyLanguageModel/​Models/​OllamaLanguageModel.swift Stabilizes request encoding; add serialized-body tests for insertion-order independence.
Sources/​AnyLanguageModel/​GenerationSchema.swift Stabilizes schema ordering; $defs keys remain unsorted and should be canonicalized.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread Sources/AnyLanguageModel/GenerationSchema.swift
Comment on lines +526 to +531
private func encodeChatParams(_ params: [String: JSONValue]) throws -> Data {
let encoder = JSONEncoder()
// Ollama reuses prompt prefixes only when their serialized bytes match.
// Dictionary iteration order must not vary between equivalent requests.
encoder.outputFormatting = [.sortedKeys]
return try encoder.encode(params)
@mattt

mattt commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Hi @akbashev. Thank you for this! It's a great catch, and the numbers from Ollama's logs make the case. Copilot pointed out two things to add before we merge:

  1. GenerationSchema.encode(to:) still encodes $defs in dictionary order, so a schema with more than one definition can still change from one request to the next. Could you sort those keys too?
  2. A test that encodes equivalent params built in different insertion orders and checks that the bytes match, so this doesn't regress.

If you'd rather not, just say so, and I'm happy to add these on top of your change.

@akbashev

Copy link
Copy Markdown
Author

@mattt good catch with $def, missed it 🤔🙂

Will do soon!

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

No unresolved blocking issues were identified.

Review effort: Lite
Findings: 1 Low severity

Open (1)
Resolved since last review (1)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants