Skip to content

chore(zenmux): sync specialized model families - #7467

Open
Alcatraz-Zhang wants to merge 8 commits into
anomalyco:devfrom
Alcatraz-Zhang:codex/zenmux-specialized-families
Open

Alcatraz-Zhang wants to merge 8 commits into
anomalyco:devfrom
Alcatraz-Zhang:codex/zenmux-specialized-families

Conversation

@Alcatraz-Zhang

@Alcatraz-Zhang Alcatraz-Zhang commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Summary

  • sync the 51 remaining live ZenMux routes across 18 specialized model families
  • add complete shared lab metadata for 38 missing underlying model identities and keep all ZenMux files base_model-based / override-only
  • remove 10 stale ZenMux routes that are absent from the current public catalog
  • preserve non-token media pricing as intentionally omitted instead of publishing invalid token costs
  • align reasoning controls with first-party or established peers; treat bfl/flux-3-video as non-reasoning because the video API exposes no reasoning request or response surface
  • correct LongCat 2.0 open-weight metadata using the official MIT-licensed Hugging Face release

This is the third independent PR in the ZenMux refresh split. It is synchronized with dev at 97816c1e8, not based on either earlier PR branch, and contains 101 changed files (under the reviewer cap of 300).

Sources

Live snapshot captured 2026-09-19:

Snapshot SHA-256:

  • page: c31783fedbf3c31fd74a5eb022395e6c6c7b51689dece6a6b8a0db3d9ae8e080
  • OpenAI: c11b2ee69e17178863b7437212ef87670d87cb8a75191831a844bec2a99feaa3
  • Anthropic: 6cae18e0e0e978c96cafa501a2ae33e3ecd821acb8747e8c1c253efe6d3604b0
  • Vertex AI: 0b80a9b9fa984e1a4d09f8b8404ac792f84305eb963c5ec3f82c80681c90603f

Open-weight identity checks used official repositories, including:

Validation

  • generator dry-run: 0 created, 0 updated, 0 removed
  • scoped structure audit: 101 files; 51 provider models; 38 added lab metadata files; no missing/extra/out-of-scope entries
  • live source audit: 51 checked, 0 mismatches
  • bun validate
  • bun run compare:migrations
  • git diff --check
  • targeted generation tests: 10 pass, 1 repository-baseline failure for 43 pre-existing open-weight entries without weight links; none are introduced by this PR

@Alcatraz-Zhang

Copy link
Copy Markdown
Contributor Author

Addressed the independent pre-review findings in a960c5b while keeping the PR Draft. Atria and Dots3 now separate official model-native capabilities/weights/licenses from ZenMux-only naming and capability limits; Llama 4 Scout now reuses the existing canonical lab identity; Ring 2.6 exposes the verified high|xhigh effort control. The source audit now also checks host temperature, tool-calling, and structured-output support, and null max_completion_tokens no longer becomes a false zero-output limit. Refreshed the live page snapshot (51-route scope unchanged) and updated its hash in the PR body. Final checks: generator 0-diff; 101-file/51-route structure audit clean; 51 source rows with zero capability/cost/modality/limit mismatches; bun validate; post-commit compare:migrations self-check; diff check. The pre-commit migration comparison cannot model the Atria base capability correction because it applies current models/ metadata to the old provider file; this is a limitation of the comparison script, not a final-tree validation failure. Targeted schema/generation tests remain 18 pass / 1 repository-baseline failure for the same 43 pre-existing missing weight links; no PR3 model is in that list.

@Alcatraz-Zhang

Copy link
Copy Markdown
Contributor Author

Before requesting automated review, synchronized upstream/dev, local dev, and origin/dev to d809a3f, then merged that latest dev into this still-Draft branch in c684bdc. The three new dev commits only add Aihubmix models and do not overlap the ZenMux scope. Revalidation after the merge is unchanged: generator 0-diff; 101-file structure audit; exact 51/51 live set; 51 source rows with zero mismatches; bun validate; compare:migrations self-check; diff check; targeted tests 18 pass / 1 unchanged repository-baseline failure.

@Alcatraz-Zhang
Alcatraz-Zhang marked this pull request as ready for review September 19, 2026 02:27
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] models/inclusionai/ling-2.6-1t.toml:2 - Check: Lab metadata must describe the model named by the file, not a different identity. Why: The description is a paste of Ring-2.6-1T (it names “Ring-2.6-1T”, calls it a thinking model with adaptive high/xhigh effort) while this lab entry is Ling-2.6-1T with reasoning = false. Consumers inherit a wrong, contradictory identity via base_model. Action: Replace the description with Ling-2.6-1T-specific text (and drop reasoning-effort claims that belong only to Ring).
  • [high] [possible mistake] providers/zenmux/models/meta/muse-spark-1.1.toml:5 - Check: Host capability overrides must be real ZenMux API deltas, not incomplete catalog flags. Why: Patch 2 forces temperature = false, tool_call = false, and structured_output = false on Muse Spark 1.1/1.2/1.3/contributor (and the same pattern on LongCat 2.0, Fugu Ultra v2 tool_call, Macaron, Agnes 2.5, Hy3, Ling 3.0 Flash/VL, Atria, etc.) while lab metadata and established peers (Meta first-party, OpenRouter, Kilo) keep tools/temperature/structured output on for these models. That systematically under-reports core capabilities if ZenMux still accepts tools. Action: Re-verify each demotion against live ZenMux request support (not only list-API booleans); remove false negatives so host files only override real gaps.
  • [medium] [possible mistake] providers/zenmux/models/xiaomi/mimo-v2.5.toml:12 - Check: [[cost.tiers]] must encode a real price change at the threshold. Why: Base and 256_000 tier rates are identical (0.12/0.232/0.00232; same pattern on mimo-v2.5-pro). A no-op tier is misleading and looks like a failed edit of the old doubled tier. Action: If pricing is flat, drop the tiers; if not, restore the correct higher tier rates with a pricing citation.
  • [medium] [possible mistake] models/alibaba/happyhorse-1.0.toml:11 - Check: open_weights must match the model’s actual release status. Why: The description calls HappyHorse 1.0 an “open-source” video model, but the lab entry sets open_weights = false with no weights link. Action: Confirm release status; set open_weights = true (+ weights/license) if open, or stop calling it open-source if closed.
  • [low] [possible mistake] providers/zenmux/models/meta/muse-spark-1.1.toml:17 - Check: Provider modality overrides should not invent inputs the underlying model does not support. Why: ZenMux lists Muse Spark 1.1 inputs as text/image/pdf/audio/video, but lab models/meta/muse-spark-1.1.toml has no audio (audio appears on 1.3). Action: Confirm ZenMux actually accepts audio on the 1.1 route; otherwise align modalities with the lab (or only override on 1.3).

@Alcatraz-Zhang
Alcatraz-Zhang marked this pull request as draft September 19, 2026 02:33
@Alcatraz-Zhang

Copy link
Copy Markdown
Contributor Author

Addressed the first automated review in 8d44954 after synchronizing upstream/dev, local dev, and origin/dev to bc21426 and merging it into this Draft branch. Corrected Ling-2.6-1T’s pasted Ring description using the official InclusionAI model card; removed systematic false capability demotions because omission from supported_parameters is not affirmative non-support; removed no-op Xiaomi pricing tiers; retained HappyHorse as closed-weight while removing the unverified open-source claim; and kept Muse Spark 1.1 audio with a leading note because the page, OpenAI list, and Anthropic list all explicitly report it. Revalidated generator idempotence, 101-file structure, exact 51/51 scope, 51 source rows with zero mismatches, bun validate, compare:migrations, and diff check. Targeted tests remain 18 pass / 1 unchanged repository-baseline failure.

@Alcatraz-Zhang
Alcatraz-Zhang marked this pull request as ready for review September 19, 2026 02:40
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/zenmux/models/sapiens-ai/agnes-2.5-flash.toml:4 - Check: Relay reasoning_options must match lab/same-surface peer controls when reasoning = true (not [] from uncertainty). Why: First-party providers/agnes/models/agnes-2.5-flash.toml exposes toggle (chat_template_kwargs.enable_thinking) plus interleaved.reasoning_content; ZenMux still authors reasoning_options = [], which means “no caller control” and understates a controlled reasoner on a multi-model relay. Same issue on agnes-2.5-pro.toml. Action: Replace [] with a leading wire comment and [[reasoning_options]] type = "toggle" (and add [interleaved] if ZenMux returns a reasoning side channel); align Pro the same way.
  • [medium] [possible mistake] models/inclusionai/ling-3.0-tiny.toml:12 - Check: New lab metadata must not contradict its own stated capabilities. Why: Description claims “native function calling” and switchable Thinking/Instant modes, but the lab file sets tool_call = false (and temperature = false) while the ZenMux host correctly adds a reasoning toggle. That leaves shared lab metadata wrong for every host that inherits it. Action: Verify InclusionAI’s real surface and set tool_call / temperature (and any other flags) to match the model card; keep host-only deltas on the ZenMux file only when ZenMux truly differs.

@Alcatraz-Zhang
Alcatraz-Zhang marked this pull request as draft September 19, 2026 02:46
@Alcatraz-Zhang

Copy link
Copy Markdown
Contributor Author

Addressed the latest automated review in fba5b09 after synchronizing upstream/dev, local dev, and origin/dev to 97816c1 and merging it into this Draft branch. Agnes 2.5 Flash/Pro now expose ZenMux reasoning.enabled toggles with reasoning_content side channels, matching the controlled reasoner surface without using empty options. Ling 3.0 Tiny lab metadata now sets temperature=true and tool_call=true, backed by InclusionAI’s official model card, which documents recommended temperature sampling plus SGLang/vLLM tool-call parsers. Revalidated generator idempotence, 101-file structure, exact 51/51 live scope, 51 source rows with zero mismatches, bun validate, compare:migrations, and diff check. Targeted generation tests remain 10 pass / 1 unchanged repository-baseline failure for 43 unrelated open-weight entries.

@Alcatraz-Zhang
Alcatraz-Zhang marked this pull request as ready for review September 19, 2026 06:40
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant