Conversation
OpenCode has no catalog entry for the gateway's system.ai.* model ids, so
every model without a configured `limit` fell back to OpenCode's
{context: 200000, output: 32000}. 1M-context models (Claude Opus/Sonnet
5.5, GLM-5.x, DeepSeek V4, Kimi K3) auto-compacted at about 180K tokens.
Apply known token limits to every managed OpenCode provider, not only
OSS. Claude Opus/Sonnet 4.6+ get a 1,000,000 context, Haiku 4.5 200,000,
and GLM-5.x, DeepSeek V4 and Kimi K3 1,048,576. GLM-4.x keeps its
200K/25K pin. Output stays at 32,000, OpenCode's own clamp.
The Claude Code [1m] suffix now shares the same claude_has_1m_context
check.
Fixes databricks#905
OpenCode auto-compacts at min(limit.input - buffer, limit.context - max(min(limit.output, 32000), buffer)). Without limit.input, 1M-context models compacted only at about 968K tokens. limit.input is only read by that check, so pin it to 90% of the context; the configurable compaction buffer then applies on top of it. Co-authored-by: Isaac <no-reply@databricks.com>
…l-back-to-opencodes-2' into feat/opencode-gateway-models-fall-back-to-opencodes-2
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #905
What
ug opencodenow gives OpenCode the real context windows of the gateway models, plus alimit.inputthat makes auto-compaction start at about 90% of the window. Token limits are applied to every managed OpenCode provider (databricks-anthropic,databricks-google,databricks-oss), not only OSS.limitafter (context/input/output)claude-opus-*/claude-sonnet-*4.6+ (incl. 5.5)claude-haiku-4-5glm-5-*deepseek-v4*,kimi-k3*glm-4-*Why
OpenCode has no catalog entry for the gateway's
system.ai.*model ids. Any model without a configuredlimitfalls back to OpenCode's hardcoded{context: 200000, output: 32000}, and that fallback can't be changed through config. As a result, sessions on 1M-context models auto-compacted at about 180K tokens instead of near the real window. The only pin ug wrote was a stale GLM-4.6{200000, 25000}entry, and it also matched GLM-5.x.Pinning
contextalone is not enough to get a sensible threshold. OpenCode 2.0.14 auto-compacts once usage reaches:bufferis the compaction buffer (20,000 by default, configurable). Withoutlimit.input, a 1M-context model compacts only at 968,000 tokens (~97%).limit.inputis read only by this check, so setting it changes when compaction starts and nothing else.How
databricks.py_MODEL_TOKEN_LIMITSis keyed by model-id substring, and the first match wins. The single"glm"entry is replaced byglm-5,deepseek-v4andkimi-k3(1,048,576),glm-4(the existing 200K/25K pin), andclaude-haiku-4-5(200K).claude_has_1m_context()covers Claude Opus/Sonnet 4.6+ by version comparison.model_token_limits()returns{1000000, 32000}for those models.agents/opencode.py_oss_model_overlaybecomes_model_overlay(model, overlay), andrender_overlayuses it for all three providers. Each Anthropic model now gets its own entry dict, where before they all shared one dict viadict.fromkeys. The per-modelheadersandtoolStreaming: falseare unchanged._model_overlaysetslimit.inputto 90% oflimit.context. This is a flat ratio and is not offset by OpenCode's buffer, so a user-configured buffer still applies on top of it. It lives in the OpenCode agent, not in the sharedmodel_token_limitstable, because it encodes OpenCode's compaction policy rather than a gateway limit.agents/claude.py:_maybe_add_1m_suffixnow usesclaude_has_1m_context, so Claude Code's[1m]suffix and OpenCode's 1M limit share one rule. Behavior is unchanged.max_tokenstomin(limit.output, OUTPUT_TOKEN_MAX = 32000), so a higher value would change nothing by default. It is also below every listed gateway cap from OpenCode: gateway models fall back to OpenCode's 200K context (Claude 5.5 and OSS are ~1M) #905 (GLM-5.x 131,072; DeepSeek V4 flash 384,000; Kimi K3 1,048,576; Claude Opus 5.5 128,000).Not changed:
_MODEL_TOKEN_LIMITSentry is enough to pin it later, sincerender_overlaynow applies limits to the Google provider too./api/2.1/unity-catalog/model-services) only returns names, so the limits stay in a table.output, which is out of scope here.Testing
uv run --frozen pytest tests/test_agent_opencode.py tests/test_databricks.py tests/test_agent_claude.py: 640 passed (after merging the latestmain). New and updated cases cover:system.ai.anddatabricks-forms), and Haiku 4.5.limit.inputat 90% of the context for every pinned model.toolStreamingstill set on every Anthropic entry.uv run --frozen pytest: 2965 passed, 40 skipped, 5 failed on the first revision of this PR, beforelimit.inputand before mergingmain. The same 5 fail onmainin this environment, so they are unrelated to this change: 2 intests/test_claude_smart_routing_v2.pyand 3 intests/test_e2e_user_agent.py. The full suite has not been re-run since.ruff check src/ tests/,ruff format --check src/ tests/, andty check src/pass.limit.inputis optional and read only by compaction.opencode servewith isolatedXDG_*/HOMEdirectories./api/modelreported the pinnedcontext/outputfor every model above, andlimit.inputwas passed through as configured. This was checked with an earlierinputformula (90% of context plus the 20K default buffer); the final flat-90% values were checked by unit tests only.Environment: macOS arm64, OpenCode 2.0.14 (
@opencode/cli).This pull request and its description were written by Isaac.