Gemma 7B
Gemma 7B is a language model with 7B parameters from Google DeepMind. It is currently free at 1 provider with up to 4K tokens of context. Its weights are open, so it can also run locally.
- Developer
- Google DeepMind
- Parameters
- 7B
- Context
- 4K tokens
- Type
- Text
- License
- Open weights
- Free since
- September 28, 2026
Where to use Gemma 7B for free
Each row is a separate free endpoint with its own limits.
| Provider | Context | Limit | Card | |
|---|---|---|---|---|
Cloudflare Workers AI @cf/google/gemma-7b-it-lora |
4K | Workers Free: 10,000 Neurons/day shared across eligible models | — | Open |
Quick start
curl https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1/chat/completions \
-H "Authorization: Bearer $CLOUDFLARE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "@cf/google/gemma-7b-it-lora",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1",
api_key="YOUR_CLOUDFLARE_API_KEY",
)
reply = client.chat.completions.create(
model="@cf/google/gemma-7b-it-lora",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1",
apiKey: process.env.CLOUDFLARE_API_KEY,
});
const reply = await client.chat.completions.create({
model: "@cf/google/gemma-7b-it-lora",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(reply.choices[0].message.content);Example for Cloudflare Workers AI. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.
Benchmarks
This model is not on the Artificial Analysis leaderboard yet. We do not publish our own estimates.
Similar free models
FAQ
Is Gemma 7B free?
Yes. As of September 28, 2026, Gemma 7B is free at Cloudflare Workers AI. Each provider has its own limits — see the table above.
What is the context window of Gemma 7B?
The largest context among free routes is 3,500 tokens. Some providers cap it lower.
How do I call Gemma 7B via API?
Use any OpenAI-compatible client with base URL https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1, model ID @cf/google/gemma-7b-it-lora, and a Cloudflare Workers AI key.
Is Gemma 7B open-weight?
Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.