#

Mixture-of-Experts

AI tools tagged Mixture-of-Experts.

31 tools · handpicked & curated
// price
DeepSeek-V4-Flash-Vision-Exp DeepSeek's experimental multimodal variant of V4-Flash for document and chart understanding
GLM-5.3-Flash Zhipu's first natively multimodal GLM-5 model, revealed after going viral under the name "Ox Alpha"
GLM-5.3 Zhipu's coding-focused update to GLM-5.2, tuned heavily on vulnerability research
DeepSeek-V4-Flash-0731 DeepSeek's MIT-licensed 284B MoE model, retrained for stronger agentic coding
Kimi K3 Moonshot's 2.8 trillion parameter open-weight model, the largest open model shipped to date
Ling-3.0-Flash Ant Group's efficient open-weight reasoning model built for production AI agents
Poolside Laguna S 2.1 Poolside's open-weight coding model that beats agentic coding rivals many times its size
Poolside Laguna XS 2.1 Poolside's smallest open-weight coding model, built to run on a single GPU
DeepSeek V3.2 DeepSeek's V3.1 update, built around a new sparse attention mechanism
Llama 5 Meta's next flagship open-weight multimodal model, built for on-device and datacenter use alike
DBRX Databricks' open-weight mixture-of-experts model
DeepSeek V3 DeepSeek's efficient mixture-of-experts model trained at a fraction of rival costs
DeepSeek V4-Pro The top open-weight model of 2026, leading on agentic coding and graduate reasoning
DeepSeek-MoE-16B DeepSeek's early fine-grained mixture-of-experts model that informed its later MoE architectures
DeepSeek-V2 DeepSeek's 2024 mixture-of-experts model that undercut rivals on API price
Hunyuan Large Tencent's open-weight mixture-of-experts model
Hunyuan-A13B Tencent's open-weight mixture-of-experts reasoning model
Kimi K2 Moonshot's original K2-generation open-weight base model
Kimi K2.6 Moonshot's open-weight model built for long-running agentic and tool-use tasks
Kimi K2.7 Code Moonshot's open-weight agentic coding model, tuned for tool-use and long-running tasks
Llama 4 Maverick Meta's flagship open-weight multimodal mixture-of-experts model
MiniMax-01 MiniMax's 2025 open-weight model with a 4M-token context window
Mixtral 8x22B Mistral's open-weight sparse mixture-of-experts model
Mixtral 8x7B Mistral's original open-weight mixture-of-experts model that popularised sparse MoE LLMs
OLMoE AI2's fully open mixture-of-experts model
PanGu-Sigma Huawei's trillion-parameter sparse mixture-of-experts model
Qwen2.5-Max Alibaba's largest proprietary MoE model, positioned against GPT-4o and DeepSeek-V3
Qwen3-235B-A22B Alibaba's flagship open-weight mixture-of-experts model, the cheapest capable frontier model
Qwen3-Coder Alibaba's open-weight agentic coding model
Samba-1 SambaNova's composition-of-experts enterprise model
Zephyr 141B Hugging Face's larger mixture-of-experts Zephyr fine-tune, distinct from the 7B release