Skip to content

Add Less is MoE (EMNLP 2026) to Mixture-of-Experts(MoE) LLM Inference - #198

Open
HectorHHZ wants to merge 1 commit into
xlite-dev:mainfrom
HectorHHZ:add-less-is-moe
Open

Add Less is MoE (EMNLP 2026) to Mixture-of-Experts(MoE) LLM Inference#198
HectorHHZ wants to merge 1 commit into
xlite-dev:mainfrom
HectorHHZ:add-less-is-moe

Conversation

@HectorHHZ

Copy link
Copy Markdown

Adding one paper to Mixture-of-Experts(MoE) LLM Inference. Disclosure: I am an author.

Why it fits. Most MoE compression drops or merges whole experts. This paper shows capability is concentrated in a tiny number of FFN intermediate dimensions — on Qwen1.5-MoE, removing 12 of 1.35M routed-FFN dimensions collapses GSM8K — and prunes at that granularity instead. At a 50% compression ratio it cuts weight memory by ~45% and improves inference throughput by 21%, so the result is an inference-side win, not only a parameter-count one. Evaluated on Qwen1.5-MoE, Qwen3-MoE and OLMoE, with vLLM and Hugging Face runtime patches in the repo.

Followed the conventions of the table: appended at the end of the MoE section in ascending date order, [[pdf]] links to the arXiv PDF, and the code link carries the usual stars badge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RheuUEHaQB6xfZuz99Jkhy
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant