#
Legacy
AI tools tagged Legacy.
abab6.5
MiniMax's proprietary flagship model prior to its open-weight pivot
ALBERT
Google's parameter-sharing variant of BERT designed for efficiency at scale
AMD Radeon Instinct MI25
AMD's first Instinct-branded datacenter accelerator
AMD Radeon Instinct MI60
Vega-based datacenter accelerator, AMD's first 7nm GPU
AMD Radeon RX Vega 64
Flagship Vega-generation consumer GPU with HBM2
Aquila
BAAI's first open bilingual (Chinese/English) foundation model
AudioLM
Google's framework for generating realistic speech and audio continuations
Baichuan 2
Baichuan's open-weight bilingual foundation model
Baichuan-13B
Baichuan's open-weight bilingual model, predating the Baichuan 2/3/4 generations
BERT
Google's 2018 bidirectional encoder that reshaped NLP research
BioGPT
Microsoft's domain-specific generative model pre-trained on biomedical literature
BlenderBot 3
Meta's open-domain chatbot research model that browses the web to stay current
BLIP
Salesforce's vision-language pretraining model for captioning and retrieval
BLOOM
The open, multilingual model built by the BigScience collaboration
BLOOMZ
An instruction-tuned variant of BLOOM trained to follow cross-lingual instructions
ByT5
Google's token-free variant of T5 that operates directly on raw bytes
ChatGLM3
Zhipu's third-generation open-weight bilingual dialogue model
Chinchilla
DeepMind's 2022 research model that reset scaling-law assumptions for LLM training
Claude 1
Anthropic's first publicly released Claude model
Claude 2
Anthropic's second-generation Claude, ahead of the Claude 2.1 refresh
Claude 2.1
Anthropic's 2023 model that pushed context length to 200K tokens
Claude 3 Haiku
Anthropic's fastest model in the original Claude 3 family
Claude 3 Opus
Anthropic's 2024 flagship, the top tier of the original Claude 3 family
Claude 3 Sonnet
The balanced mid-tier of Anthropic's original Claude 3 lineup
Claude Instant
Anthropic's original fast, low-cost model tier
CLIP
OpenAI's foundational image-text contrastive model that underlies much of modern multimodal AI
Codex
OpenAI's original code-completion model that powered early GitHub Copilot
Command Light
Cohere's smallest, fastest Command tier
DALL-E
OpenAI's original 2021 text-to-image model
DALL-E 2
OpenAI's second image generator, the first to bring photorealistic text-to-image to the mainstream
DALL-E 3
OpenAI's image generation model integrated directly into ChatGPT
DeepFloyd IF
A pixel-space (non-latent) text-to-image model developed by the Stability-backed DeepFloyd lab
DeepSeek LLM 67B
DeepSeek's first dense open-weight foundation model
DeepSeek-Coder-33B
DeepSeek's earlier flagship code model, preceding DeepSeek-Coder-V2
DeepSeek-MoE-16B
DeepSeek's early fine-grained mixture-of-experts model that informed its later MoE architectures
DeepSeek-V2
DeepSeek's 2024 mixture-of-experts model that undercut rivals on API price
DeepSeek-VL
DeepSeek's first vision-language model, preceding DeepSeek-VL2
DistilBERT
Hugging Face's distilled, 40%-smaller version of BERT that retains most of its performance
Dolly 2.0
Databricks' first fully open, commercially-usable instruction-following model
ELECTRA
Google's more sample-efficient pretraining approach using a replaced-token-detection objective
ELMo
AI2's deep contextualized word representation model that predated the transformer era
Emu
Meta's photorealistic text-to-image model behind Meta AI's image features
Emu Video
Meta's text-to-video model built on the Emu image generator
ERNIE 4.0
Baidu's flagship model prior to the 4.5 generation
EXAONE 3.0
LG's earlier EXAONE generation, preceding EXAONE 3.5
Falcon 180B
TII's open-weight model that led open leaderboards on release
Falcon 40B
TII's earlier open-weight model that led leaderboards before Falcon 180B
Falcon 7B
TII's compact Falcon tier that helped popularize the original Falcon release
Flan-T5
Google's instruction-tuned, open-weight successor to T5
Galactica
Meta's 2022 model trained on scientific literature, pulled days after launch
Gemini 1.0 Pro
The mid-tier of Google's first Gemini generation
Gemini 1.0 Ultra
Google's first Gemini-generation flagship, launched at the start of 2024
Gemini 1.5 Pro
Google's 2024 model that introduced the 1M-token context window
Gemma
Google's original open-weight Gemma release, built from the same research as Gemini
Google TPU v1
Google's first-generation inference-only TPU, never sold externally
Gopher
DeepMind's 280B-parameter research model that informed the Chinchilla scaling laws
GPT-1
OpenAI's original 2018 paper model that introduced the GPT architecture
GPT-2
OpenAI's 2019 model, once withheld from release over misuse concerns
GPT-3
The 2020 model that first showed large-scale few-shot learning was possible
GPT-3 Davinci
The largest of the original GPT-3 model sizes
GPT-3.5 Turbo
The model that powered ChatGPT's original public launch
GPT-4
The 2023 release that reset expectations for what LLMs could do
GPT-J
EleutherAI's early open-weight GPT-3-style model
GPT-Neo 2.7B
EleutherAI's early GPT-3-style open model, a predecessor to GPT-J and GPT-NeoX
GPT-NeoX-20B
EleutherAI's open-weight model, one of the largest public checkpoints of its era
Granite 13B
IBM's earlier enterprise foundation model, preceding the Granite 3 generation
Grok-1
xAI's first model, later open-weighted under Apache 2.0
Grok-1.5
xAI's transitional model that added long-context reasoning ahead of Grok 2
Grok-1.5V
xAI's first vision-capable Grok model, announced as a preview ahead of wider release
Ideogram 1.0
Ideogram's first model, notable for unusually reliable in-image text rendering
Imagen
Google's original text-to-image diffusion model, introduced alongside Parti as a research preview
Imagen 2
Google's second-generation text-to-image model with improved photorealism and text rendering
Inflection-1
Inflection's original model that first powered the Pi assistant
InstructGPT
OpenAI's 2022 model that introduced RLHF instruction-tuning at scale
Jina Embeddings v2
Jina AI's embedding model, notable for an 8k-token context window at launch
Jurassic-1
AI21's first large language model, a GPT-3 era contemporary
Jurassic-2
AI21's earlier proprietary large language model line
Jurassic-2 Mid
The mid-sized tier of AI21's Jurassic-2 family, preceding the Jamba architecture switch
LaMDA
Google's 2021 conversational model that underpinned the original Bard
Leonardo Diffusion XL
Leonardo's earlier fine-tuned Stable Diffusion XL base model
LLaMA
Meta's original 2023 open-weight release that kicked off the open LLM boom
Llama 2
Meta's 2023 open-weight release that jump-started the open LLM ecosystem
Llama 3
Meta's 2024 open-weight release that closed most of the gap to closed models
Llama Guard 2
Meta's earlier content-safety classifier model, preceding Llama Guard 3
Luma Dream Machine 1.0
Luma's first publicly available text/image-to-video model
Megatron-Turing NLG
NVIDIA and Microsoft's 530B-parameter research model, among the largest dense LLMs of its time
Mistral Large 2
Mistral's 2024 flagship model, the generation preceding Mistral Large 3
MPT-30B
MosaicML's open-weight commercial-use model, later folded into Databricks
mT5
Google's multilingual variant of T5, covering 101 languages
Muse
Google's masked-transformer text-to-image model, faster than comparable diffusion models
MusicLM
Google's early text-to-music generation model
Nomic Embed Text v1
Nomic's first fully open-source, open-weight, open-data long-context embedding model
NVIDIA GeForce GTX 980 Ti
Flagship Maxwell-era consumer GPU
NVIDIA Tesla K80
Dual-GPU Kepler-era compute accelerator
NVIDIA Tesla M60
Dual-GPU Maxwell-era virtualization card
NVIDIA Titan Xp
Prosumer flagship of the Pascal generation
o1-mini
OpenAI's first compact reasoning model, tuned for coding and math
o1-preview
Public preview of OpenAI's first reasoning model, ahead of the full o1 release
OLMo 7B
AI2's first fully open model release, with open weights, data, and training code
OPT-175B
Meta's 2022 open-weight model released to mirror GPT-3's scale
PaLM
Google's 2022 540B-parameter model that set the stage for Gemini
PaLM 2
Google's 2023 large language model that powered the original Bard
PanGu-Alpha
Huawei's first large-scale Chinese autoregressive language model
Parti
Google's autoregressive text-to-image model, an alternative approach to diffusion
Phi-2
Microsoft's earlier 2.7B model demonstrating outsized small-model performance
Qwen-7B
Alibaba's original Qwen release that started the model family
Qwen-Audio
Alibaba's first audio-language model for speech and sound understanding
Qwen-VL
Alibaba's first vision-language model in the Qwen family
Qwen1.5-110B
Alibaba's largest dense model in the Qwen1.5 generation
Qwen2-72B
Alibaba's second-generation flagship dense open-weight model
QwQ-32B-Preview
Alibaba's first public preview of its QwQ reasoning line, ahead of the full QwQ-32B release
Recraft V2
Recraft's earlier image and vector generation model, preceding V3
RoBERTa
Meta's robustly-optimized retraining of BERT that improved on it across benchmarks
Runway Gen-1
Runway's first video-to-video generation model
Runway Gen-2
Runway's first text-to-video model, preceding Gen-3 Alpha
Runway Gen-3 Alpha
Runway's 2024 text-to-video generation model
Samsung Gauss
Samsung's original in-house foundation model, ahead of Gauss 2
SantaCoder
BigCode's small multilingual code model preceding StarCoder
SenseChat-3
SenseTime's third-generation SenseChat model, preceding SenseChat-5
SenseNova 5.0
SenseTime's flagship multimodal SenseNova model, ahead of the 5.5 update
Skywork-13B
Kunlun Tech's open-weight bilingual foundation model
SmolLM
Hugging Face's original family of small, fully open language models
Solar 10.7B
Upstage's depth-upscaled open model that preceded the Solar Pro line
Sora
OpenAI's original text-to-video model, notable for minute-long coherent generations at launch
Sparrow
DeepMind's dialogue agent research model trained to be more helpful and less harmful
StableLM 3B
Stability AI's early compact open-weight language model, preceding StableLM 2
StarCoder
The original BigCode model that StarCoder2 later succeeded
Step-1
StepFun's earlier flagship model that preceded Step-2
T5
Google's text-to-text transfer transformer, an early unifying NLP framework
Tacotron 2
Google's neural text-to-speech model pairing a spectrogram predictor with a WaveNet vocoder
text-davinci-003
The last and most capable of OpenAI's original GPT-3.5 completion models
text-embedding-ada-002
OpenAI's long-serving default embedding model before the v3 family
Titan Text
Amazon's original in-house foundation model line for Bedrock
Titan Text Express
Amazon's mid-tier Titan model for general text generation, between Lite and the base Titan Text
Turing-NLG
Microsoft's 17B-parameter model that was among the largest published language models of its time
UL2
Google's unified pretraining framework blending multiple denoising objectives
Veo
Google's original Veo text-to-video model, previewed at I/O 2024 ahead of Veo 3
Vicuna-13B
LMSYS's influential early open chat model fine-tuned from Llama on ShareGPT conversations
Vicuna-33B
The largest Vicuna release, LMSYS's ShareGPT fine-tune of Llama
Voyage-2
Voyage AI's general-purpose embedding model, preceding Voyage-3
Wav2Vec 2.0
Meta's self-supervised speech representation model that underpins many ASR systems
WaveNet
DeepMind's raw-audio generative model that set the foundation for modern neural TTS
WizardCoder
Microsoft's code-specialised fine-tune built with Evol-Instruct data generation
XGen-7B
Salesforce's open-weight long-context research model
XLNet
A permutation-based autoregressive pretraining model that outperformed BERT on many tasks