#
Low-Latency
AI tools tagged Low-Latency.
Gemini 3.7 Flash
Google's successor to Gemini 3.6 Flash, shipped three weeks later with cheaper output pricing
Gemini 3.6 Flash
Google's new default Flash model, cheaper and faster on coding and agentic tasks than 3.5 Flash
Gemini 3 Flash
Google's Flash-tier model bringing Gemini 3 Pro reasoning to lower cost and latency
Gemini 3.5 Flash
Google's default Flash model, running about four times faster than rival frontier models
Amazon Nova Lite
Amazon's low-cost multimodal Bedrock model tier
Claude 3 Haiku
Anthropic's fastest model in the original Claude 3 family
Claude 3.5 Haiku
Anthropic's fast, affordable model matching the prior Claude 3 Opus on some benchmarks
Claude Haiku 4.5
Anthropic's fastest and cheapest current-generation model
Claude Instant
Anthropic's original fast, low-cost model tier
Command Light
Cohere's smallest, fastest Command tier
Command R7B
Cohere's smallest current-generation Command model
Gemini 2.0 Flash
Google's low-latency 2.0-generation multimodal model
Gemini 2.5 Flash
Google's low-latency, cost-efficient Gemini tier
GPT-4o mini
OpenAI's small, cost-efficient multimodal model
Grok 4 Fast
xAI's low-latency tier of Grok 4 for high-throughput use
Hunyuan Turbo
Tencent's low-latency Hunyuan tier for production workloads
o3-mini
OpenAI's small, fast reasoning model for STEM tasks
Sonic
Cartesia's low-latency real-time text-to-speech model
text-embedding-3-small
OpenAI's smaller, cheaper text embedding model