Add Less is MoE (EMNLP 2026) to Mixture-of-Experts(MoE) LLM Inference - #198
Open
HectorHHZ wants to merge 1 commit into
Open
Add Less is MoE (EMNLP 2026) to Mixture-of-Experts(MoE) LLM Inference#198HectorHHZ wants to merge 1 commit into
HectorHHZ wants to merge 1 commit into
Conversation
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RheuUEHaQB6xfZuz99Jkhy
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adding one paper to Mixture-of-Experts(MoE) LLM Inference. Disclosure: I am an author.
Why it fits. Most MoE compression drops or merges whole experts. This paper shows capability is concentrated in a tiny number of FFN intermediate dimensions — on Qwen1.5-MoE, removing 12 of 1.35M routed-FFN dimensions collapses GSM8K — and prunes at that granularity instead. At a 50% compression ratio it cuts weight memory by ~45% and improves inference throughput by 21%, so the result is an inference-side win, not only a parameter-count one. Evaluated on Qwen1.5-MoE, Qwen3-MoE and OLMoE, with vLLM and Hugging Face runtime patches in the repo.
Followed the conventions of the table: appended at the end of the MoE section in ascending date order,
[[pdf]]links to the arXiv PDF, and the code link carries the usual stars badge.