Llama 3.1 Nemotron 70B
Llama 3.1 Nemotron 70B is a language model with 70B parameters from NVIDIA. It is currently free at 2 providers with up to 16K tokens of context. Its weights are open, so it can also run locally.
- Developer
- NVIDIA
- Parameters
- 70B
- Context
- 16K tokens
- Type
- Text
- License
- Open weights
Where to use Llama 3.1 Nemotron 70B for free
Each row is a separate free endpoint with its own limits.
| Provider | Context | Limit | Card | |
|---|---|---|---|---|
Together AI nim/nvidia/llama-3.1-nemotron-70b-instruct |
16K | Fair use (depends on server load) | — | Open |
NVIDIA Build nvidia/llama-3.1-nemotron-70b-instruct |
— | Free serverless inference for development; limits and model availability may change. | No | Open |
Quick start
curl https://api.together.xyz/v1/chat/completions \
-H "Authorization: Bearer $TOGETHER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nim/nvidia/llama-3.1-nemotron-70b-instruct",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.together.xyz/v1",
api_key="YOUR_TOGETHER_API_KEY",
)
reply = client.chat.completions.create(
model="nim/nvidia/llama-3.1-nemotron-70b-instruct",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.together.xyz/v1",
apiKey: process.env.TOGETHER_API_KEY,
});
const reply = await client.chat.completions.create({
model: "nim/nvidia/llama-3.1-nemotron-70b-instruct",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(reply.choices[0].message.content);Example for Together AI. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.
Benchmarks
Source: Artificial Analysis · Llama 3.1 Nemotron Instruct 70B
Similar free models
FAQ
Is Llama 3.1 Nemotron 70B free?
Yes. As of September 26, 2026, Llama 3.1 Nemotron 70B is free at Together AI, NVIDIA Build. Each provider has its own limits — see the table above.
What is the context window of Llama 3.1 Nemotron 70B?
The largest context among free routes is 16,384 tokens. Some providers cap it lower.
How do I call Llama 3.1 Nemotron 70B via API?
Use any OpenAI-compatible client with base URL https://api.together.xyz/v1, model ID nim/nvidia/llama-3.1-nemotron-70b-instruct, and a Together AI key.
How good is Llama 3.1 Nemotron 70B?
Artificial Analysis Intelligence Index: 6.9, GPQA Diamond: 47%. For comparison, the strongest free model in the catalog right now is GLM 5.3.
Is Llama 3.1 Nemotron 70B open-weight?
Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.