Nemotron 5
NVIDIA's open-weight model tuned for efficient inference on its own hardware
Nemotron 5 continues NVIDIA’s line of open-weight models built less to top leaderboards and more to showcase how well models run on the company’s own GPUs and inference stack. It’s a 70 billion parameter dense model with a 128,000 token context window, distributed under the NVIDIA Open Model License and packaged for one-click deployment through NVIDIA NIM microservices as well as plain Hugging Face downloads.
NVIDIA has used the Nemotron line since Nemotron-4 340B to demonstrate techniques like pruning and distillation that shrink models without gutting quality, and Nemotron 5 keeps that focus. The 70B size was chosen specifically to fit on a single H100 or a pair of consumer GPUs at reasonable quantization. NVIDIA’s real pitch here is the hardware and software stack underneath the model, not the model itself.