LLM Perks

Llama 3.3 Nemotron Super 49B

Free right now Mid-size
Tool useOpen weights

Llama 3.3 Nemotron Super 49B is a language model with 49B parameters from NVIDIA. It is currently free at 1 provider with up to 16K tokens of context. The model supports tool calling. Its weights are open, so it can also run locally.

Developer
NVIDIA
Parameters
49B
Context
16K tokens
Type
Text
License
Open weights

Where to use Llama 3.3 Nemotron Super 49B for free

Each row is a separate free endpoint with its own limits.

ProviderContextLimitCard
Together AI
nim/nvidia/llama-3.3-nemotron-super-49b-v1
16K Fair use (depends on server load) — Open

Quick start

curl https://api.together.xyz/v1/chat/completions \
  -H "Authorization: Bearer $TOGETHER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nim/nvidia/llama-3.3-nemotron-super-49b-v1",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Example for Together AI. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.

Benchmarks

Intelligence Index
8.9
GPQA Diamond
64.3%
Humanity's Last Exam
5.6%
Terminal-Bench Hard
0.0%
τ²-Bench
26.9%
Long context (LCR)
18.3%

Source: Artificial Analysis · Llama 3.3 Nemotron Super 49B v1 (Reasoning)

Similar free models

FAQ

Is Llama 3.3 Nemotron Super 49B free?

Yes. As of September 26, 2026, Llama 3.3 Nemotron Super 49B is free at Together AI. Each provider has its own limits — see the table above.

What is the context window of Llama 3.3 Nemotron Super 49B?

The largest context among free routes is 16,384 tokens. Some providers cap it lower.

How do I call Llama 3.3 Nemotron Super 49B via API?

Use any OpenAI-compatible client with base URL https://api.together.xyz/v1, model ID nim/nvidia/llama-3.3-nemotron-super-49b-v1, and a Together AI key.

How good is Llama 3.3 Nemotron Super 49B?

Artificial Analysis Intelligence Index: 8.9, GPQA Diamond: 64%. For comparison, the strongest free model in the catalog right now is GLM 5.3.

Is Llama 3.3 Nemotron Super 49B open-weight?

Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.