Falcon Mamba 7B

TII's open-weight model built on the state-space Mamba architecture

Free Language
Visit Product Page →
8192 context tokens
7B parameters
TII Falcon Mamba 7B License 1.0 license
Aug 2024 released

Falcon Mamba 7B, released by the Technology Innovation Institute in August 2024, was the first strong open 7B language model built entirely on the Mamba state-space architecture instead of transformer attention. Because it has no attention mechanism, memory use and inference speed stay constant as generated sequences get longer, unlike transformers whose compute grows with context length. TII trained it in stages that pushed the effective context handling from 2,048 up to 8,192 tokens. In benchmark comparisons it matched or beat transformer models of similar size, including Mistral 7B and Meta’s Llama 3.1 8B, and outperformed other non-transformer models such as RecurrentGemma 9B and RWKV-v6. TII distributes the weights on Hugging Face under its own Falcon Mamba license, aimed at researchers and developers who want an alternative to attention-based architectures for long-sequence workloads.