#
Mixture-of-Experts
AI tools tagged Mixture-of-Experts.
DeepSeek-V4-Flash-Vision-Exp
DeepSeek's experimental multimodal variant of V4-Flash for document and chart understanding
GLM-5.3-Flash
Zhipu's first natively multimodal GLM-5 model, revealed after going viral under the name "Ox Alpha"
GLM-5.3
Zhipu's coding-focused update to GLM-5.2, tuned heavily on vulnerability research
DeepSeek-V4-Flash-0731
DeepSeek's MIT-licensed 284B MoE model, retrained for stronger agentic coding
Kimi K3
Moonshot's 2.8 trillion parameter open-weight model, the largest open model shipped to date
Ling-3.0-Flash
Ant Group's efficient open-weight reasoning model built for production AI agents
Poolside Laguna S 2.1
Poolside's open-weight coding model that beats agentic coding rivals many times its size
Poolside Laguna XS 2.1
Poolside's smallest open-weight coding model, built to run on a single GPU
DeepSeek V3.2
DeepSeek's V3.1 update, built around a new sparse attention mechanism
Llama 5
Meta's next flagship open-weight multimodal model, built for on-device and datacenter use alike
DBRX
Databricks' open-weight mixture-of-experts model
DeepSeek V3
DeepSeek's efficient mixture-of-experts model trained at a fraction of rival costs
DeepSeek V4-Pro
The top open-weight model of 2026, leading on agentic coding and graduate reasoning
DeepSeek-MoE-16B
DeepSeek's early fine-grained mixture-of-experts model that informed its later MoE architectures
DeepSeek-V2
DeepSeek's 2024 mixture-of-experts model that undercut rivals on API price
Hunyuan Large
Tencent's open-weight mixture-of-experts model
Hunyuan-A13B
Tencent's open-weight mixture-of-experts reasoning model
Kimi K2
Moonshot's original K2-generation open-weight base model
Kimi K2.6
Moonshot's open-weight model built for long-running agentic and tool-use tasks
Kimi K2.7 Code
Moonshot's open-weight agentic coding model, tuned for tool-use and long-running tasks
Llama 4 Maverick
Meta's flagship open-weight multimodal mixture-of-experts model
MiniMax-01
MiniMax's 2025 open-weight model with a 4M-token context window
Mixtral 8x22B
Mistral's open-weight sparse mixture-of-experts model
Mixtral 8x7B
Mistral's original open-weight mixture-of-experts model that popularised sparse MoE LLMs
OLMoE
AI2's fully open mixture-of-experts model
PanGu-Sigma
Huawei's trillion-parameter sparse mixture-of-experts model
Qwen2.5-Max
Alibaba's largest proprietary MoE model, positioned against GPT-4o and DeepSeek-V3
Qwen3-235B-A22B
Alibaba's flagship open-weight mixture-of-experts model, the cheapest capable frontier model
Qwen3-Coder
Alibaba's open-weight agentic coding model
Samba-1
SambaNova's composition-of-experts enterprise model
Zephyr 141B
Hugging Face's larger mixture-of-experts Zephyr fine-tune, distinct from the 7B release