Skip to content

Pin OpenCode context limits for gateway Claude and OSS models - #907

Open
mats16 wants to merge 7 commits into
databricks:mainfrom
mats16:feat/opencode-gateway-models-fall-back-to-opencodes-2
Open

mats16 wants to merge 7 commits into
databricks:mainfrom
mats16:feat/opencode-gateway-models-fall-back-to-opencodes-2

Conversation

@mats16

@mats16 mats16 commented Sep 30, 2026 •

Copy link
Copy Markdown

Fixes #905

What

ug opencode now gives OpenCode the real context windows of the gateway models, plus a limit.input that makes auto-compaction start at about 90% of the window. Token limits are applied to every managed OpenCode provider (databricks-anthropic, databricks-google, databricks-oss), not only OSS.

Model limit after (context / input / output) Auto-compaction before Auto-compaction after (default buffer)
claude-opus-* / claude-sonnet-* 4.6+ (incl. 5.5) 1,000,000 / 900,000 / 32,000 180,000 (OpenCode fallback) 880,000
claude-haiku-4-5 200,000 / 180,000 / 32,000 180,000 (fallback) 160,000
glm-5-* 1,048,576 / 943,718 / 32,000 175,000 (ug GLM-4.6 pin) 923,718
deepseek-v4*, kimi-k3* 1,048,576 / 943,718 / 32,000 180,000 (fallback) 923,718
glm-4-* 200,000 / 180,000 / 25,000 175,000 (ug pin) 160,000
Gemini and other unverified models not pinned 180,000 (fallback) unchanged

Why

OpenCode has no catalog entry for the gateway's system.ai.* model ids. Any model without a configured limit falls back to OpenCode's hardcoded {context: 200000, output: 32000}, and that fallback can't be changed through config. As a result, sessions on 1M-context models auto-compacted at about 180K tokens instead of near the real window. The only pin ug wrote was a stale GLM-4.6 {200000, 25000} entry, and it also matched GLM-5.x.

Pinning context alone is not enough to get a sensible threshold. OpenCode 2.0.14 auto-compacts once usage reaches:

min(limit.input - buffer, limit.context - max(min(limit.output, 32000), buffer))

buffer is the compaction buffer (20,000 by default, configurable). Without limit.input, a 1M-context model compacts only at 968,000 tokens (~97%). limit.input is read only by this check, so setting it changes when compaction starts and nothing else.

How

  • databricks.py
    • _MODEL_TOKEN_LIMITS is keyed by model-id substring, and the first match wins. The single "glm" entry is replaced by glm-5, deepseek-v4 and kimi-k3 (1,048,576), glm-4 (the existing 200K/25K pin), and claude-haiku-4-5 (200K).
    • New claude_has_1m_context() covers Claude Opus/Sonnet 4.6+ by version comparison. model_token_limits() returns {1000000, 32000} for those models.
  • agents/opencode.py
    • _oss_model_overlay becomes _model_overlay(model, overlay), and render_overlay uses it for all three providers. Each Anthropic model now gets its own entry dict, where before they all shared one dict via dict.fromkeys. The per-model headers and toolStreaming: false are unchanged.
    • _model_overlay sets limit.input to 90% of limit.context. This is a flat ratio and is not offset by OpenCode's buffer, so a user-configured buffer still applies on top of it. It lives in the OpenCode agent, not in the shared model_token_limits table, because it encodes OpenCode's compaction policy rather than a gateway limit.
  • agents/claude.py: _maybe_add_1m_suffix now uses claude_has_1m_context, so Claude Code's [1m] suffix and OpenCode's 1M limit share one rule. Behavior is unchanged.
  • Output is 32,000 because OpenCode clamps max_tokens to min(limit.output, OUTPUT_TOKEN_MAX = 32000), so a higher value would change nothing by default. It is also below every listed gateway cap from OpenCode: gateway models fall back to OpenCode's 200K context (Claude 5.5 and OSS are ~1M) #905 (GLM-5.x 131,072; DeepSeek V4 flash 384,000; Kimi K3 1,048,576; Claude Opus 5.5 128,000).

Not changed:

  • Gemini stays unpinned because its gateway limits have not been verified (OpenCode: gateway models fall back to OpenCode's 200K context (Claude 5.5 and OSS are ~1M) #905 notes the same). Adding a _MODEL_TOKEN_LIMITS entry is enough to pin it later, since render_overlay now applies limits to the Google provider too.
  • Model discovery (/api/2.1/unity-catalog/model-services) only returns names, so the limits stay in a table.
  • 200K-context models compact at 80% with the default buffer. Raising that would need a lower output, which is out of scope here.

Testing

  • uv run --frozen pytest tests/test_agent_opencode.py tests/test_databricks.py tests/test_agent_claude.py: 640 passed (after merging the latest main). New and updated cases cover:
    • Claude Opus/Sonnet 5.5 and 4.6 (both system.ai. and databricks- forms), and Haiku 4.5.
    • GLM-5.x, DeepSeek V4 flash and Kimi K3.
    • limit.input at 90% of the context for every pinned model.
    • The GLM-4.x pin kept.
    • Models without known limits (Sonnet 4.5, Fable, Kimi K2, Gemini) left unpinned.
    • Per-model toolStreaming still set on every Anthropic entry.
  • uv run --frozen pytest: 2965 passed, 40 skipped, 5 failed on the first revision of this PR, before limit.input and before merging main. The same 5 fail on main in this environment, so they are unrelated to this change: 2 in tests/test_claude_smart_routing_v2.py and 3 in tests/test_e2e_user_agent.py. The full suite has not been re-run since.
  • ruff check src/ tests/, ruff format --check src/ tests/, and ty check src/ pass.
  • Manual check with OpenCode 2.0.14:
    • Read the auto-compaction check and the model fallback from the installed binary to confirm the formula above, and that limit.input is optional and read only by compaction.
    • Rendered the overlay and ran opencode serve with isolated XDG_*/HOME directories. /api/model reported the pinned context/output for every model above, and limit.input was passed through as configured. This was checked with an earlier input formula (90% of context plus the 20K default buffer); the final flat-90% values were checked by unit tests only.
    • No model requests were sent.

Environment: macOS arm64, OpenCode 2.0.14 (@opencode/cli).

This pull request and its description were written by Isaac.

mats16 and others added 7 commits September 30, 2026 17:49
OpenCode has no catalog entry for the gateway's system.ai.* model ids, so
every model without a configured `limit` fell back to OpenCode's
{context: 200000, output: 32000}. 1M-context models (Claude Opus/Sonnet
5.5, GLM-5.x, DeepSeek V4, Kimi K3) auto-compacted at about 180K tokens.

Apply known token limits to every managed OpenCode provider, not only
OSS. Claude Opus/Sonnet 4.6+ get a 1,000,000 context, Haiku 4.5 200,000,
and GLM-5.x, DeepSeek V4 and Kimi K3 1,048,576. GLM-4.x keeps its
200K/25K pin. Output stays at 32,000, OpenCode's own clamp.

The Claude Code [1m] suffix now shares the same claude_has_1m_context
check.

Fixes databricks#905
OpenCode auto-compacts at min(limit.input - buffer, limit.context -
max(min(limit.output, 32000), buffer)). Without limit.input, 1M-context
models compacted only at about 968K tokens. limit.input is only read by
that check, so pin it to 90% of the context; the configurable compaction
buffer then applies on top of it.

Co-authored-by: Isaac <no-reply@databricks.com>
…l-back-to-opencodes-2' into feat/opencode-gateway-models-fall-back-to-opencodes-2
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenCode: gateway models fall back to OpenCode's 200K context (Claude 5.5 and OSS are ~1M)

1 participant