Llama 3.1 Nemotron Ultra 253B
Llama 3.1 Nemotron Ultra 253B is a language model with 253B parameters from NVIDIA. It is currently free at 1 provider with up to — tokens of context. The model supports tool calling. Its weights are open, so it can also run locally.
- Developer
- NVIDIA
- Parameters
- 253B
- Type
- Text
- License
- Open weights
Where to use Llama 3.1 Nemotron Ultra 253B for free
Each row is a separate free endpoint with its own limits.
| Provider | Context | Limit | Card | |
|---|---|---|---|---|
NVIDIA Build nvidia/llama-3.1-nemotron-ultra-253b-v1 |
— | Free serverless inference for development; limits and model availability may change. | No | Open |
Quick start
curl https://integrate.api.nvidia.com/v1/chat/completions \
-H "Authorization: Bearer $NVIDIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/llama-3.1-nemotron-ultra-253b-v1",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://integrate.api.nvidia.com/v1",
api_key="YOUR_NVIDIA_API_KEY",
)
reply = client.chat.completions.create(
model="nvidia/llama-3.1-nemotron-ultra-253b-v1",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://integrate.api.nvidia.com/v1",
apiKey: process.env.NVIDIA_API_KEY,
});
const reply = await client.chat.completions.create({
model: "nvidia/llama-3.1-nemotron-ultra-253b-v1",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(reply.choices[0].message.content);Example for NVIDIA Build. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.
Benchmarks
Source: Artificial Analysis · Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)
Similar free models
FAQ
Is Llama 3.1 Nemotron Ultra 253B free?
Yes. As of September 26, 2026, Llama 3.1 Nemotron Ultra 253B is free at NVIDIA Build. Each provider has its own limits — see the table above.
How do I call Llama 3.1 Nemotron Ultra 253B via API?
Use any OpenAI-compatible client with base URL https://integrate.api.nvidia.com/v1, model ID nvidia/llama-3.1-nemotron-ultra-253b-v1, and a NVIDIA Build key.
How good is Llama 3.1 Nemotron Ultra 253B?
Artificial Analysis Intelligence Index: 7.5, GPQA Diamond: 73%. For comparison, the strongest free model in the catalog right now is GLM 5.3.
Is Llama 3.1 Nemotron Ultra 253B open-weight?
Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.