# LLM Perks > A live catalog of language models with free API access — providers, limits, context windows, benchmarks — plus free AI software and AI credits. Snapshot: 2026-09-26T20:55:10.329Z Free models: 214 (294 endpoints) · Providers online: 15/15 Benchmarks: Artificial Analysis (2026-09-24), https://artificialanalysis.ai/ ## Sections - [Free model catalog](https://llmperks.com) - [Model families](https://llmperks.com/models) - [Provider comparison](https://llmperks.com/providers) - [New free models](https://llmperks.com/new) - [Free AI software](https://llmperks.com/software) - [Free AI credits](https://llmperks.com/credits) - [AI subscription ROI](https://llmperks.com/subscriptions) - [Infrastructure](https://llmperks.com/infrastructure) - [Blog](https://llmperks.com/blog) - [Full model dump](https://llmperks.com/llms-full.txt) ## Best free models right now - [GLM 5.3](https://llmperks.com/glm/glm-5-3): AA index 44.8, 1048576 ctx, free at together, nvidia - [Mimo V2.6 Pro](https://llmperks.com/models/mimo-v2-6-pro): AA index 46.3, 1000000 ctx, free at unorouter - [Kimi K3](https://llmperks.com/kimi/kimi-k3): AA index 43.6, 1000000 ctx, free at nvidia, unorouter - [GLM 5.3 Flash](https://llmperks.com/glm/glm-5-3-flash): AA index 41.8, free at orcarouter, nvidia - [DeepSeek V4.1 Flash](https://llmperks.com/deepseek/deepseek-v4-1-flash): AA index 39.5, 1048576 ctx, free at nvidia, unorouter - [Qwen 3.8 Flash Next](https://llmperks.com/qwen/qwen3-8-flash-next): AA index 39.8, 1000000 ctx, free at unorouter - [DeepSeek V4 Flash](https://llmperks.com/deepseek/deepseek-v4-flash): AA index 34.3, 1000000 ctx, free at orcarouter, unorouter - [Qwen3.8 27B](https://llmperks.com/qwen/qwen3-8-27b): AA index 33.7, 262144 ctx, free at openrouter, groq, unorouter - [Gemini 3.6 Flash](https://llmperks.com/gemini/gemini-3-6-flash): AA index 34.0, 1000000 ctx, free at unorouter - [GLM 5.2](https://llmperks.com/glm/glm-5-2): AA index 33.7, 1048576 ctx, free at together - [Qwen 3.8 27B](https://llmperks.com/qwen/qwen-3-8-27b): AA index 33.7, free at cerebras - [Inkling Small](https://llmperks.com/inkling/inkling-small): AA index 27.8, 1048576 ctx, free at openrouter - [Solar Pro 4](https://llmperks.com/solar/solar-pro4): AA index 28.2, 524288 ctx, free at nous - [Nemotron 3 Ultra](https://llmperks.com/nemotron/nemotron-3-ultra): AA index 22.9, 1048576 ctx, free at openrouter, nvidia, unorouter - [Mimo V2.5](https://llmperks.com/models/mimo-v2-5): AA index 25.2, 1000000 ctx, free at unorouter - [Inkling](https://llmperks.com/inkling/inkling): AA index 25.0, 1048576 ctx, free at openrouter - [Kimi K2.6](https://llmperks.com/kimi/kimi-k2-6): AA index 27.0, free at nvidia - [Ling 3.0 Flash Fin](https://llmperks.com/ling/ling-3-0-flash-fin): AA index 22.6, 262144 ctx, free at openrouter, nous, unorouter - [Hy3](https://llmperks.com/hunyuan/hy3): AA index 25.3, free at orcarouter - [Gemini 3.5 Flash Lite](https://llmperks.com/gemini/gemini-3-5-flash-lite): AA index 22.2, 1000000 ctx, free at unorouter - [Gemma 4 31B](https://llmperks.com/gemma/gemma-4-31b): AA index 19.0, 262144 ctx, free at openrouter, together, nvidia - [GLM 4.7](https://llmperks.com/glm/glm-4-7): AA index 22.2, 202752 ctx, free at together - [Step 3.7 Flash](https://llmperks.com/step/step-3-7-flash): AA index 19.5, 262144 ctx, free at nous, unorouter - [LongCat 2.0](https://llmperks.com/longcat/longcat-2-0): AA index 19.1, 1048756 ctx, free at nous - [Gemma 4 26B A4B](https://llmperks.com/gemma/gemma-4-26b): AA index 16.7, 262144 ctx, free at openrouter, together, unorouter ## Use cases - [Free models for AI agents](https://llmperks.com/for/ai-agents) - [Free models for coding](https://llmperks.com/for/coding) - [Free reasoning models](https://llmperks.com/for/reasoning) - [Free long-context models](https://llmperks.com/for/long-context) - [Free vision models](https://llmperks.com/for/vision) - [Free image generation models](https://llmperks.com/for/image) - [Free video generation models](https://llmperks.com/for/video) - [Free music and audio models](https://llmperks.com/for/music) ## Providers - [OpenRouter](https://llmperks.com/openrouter): Models with the :free suffix allow 50 requests per day, or 1,000 per day after a one-time $10 top-up (20 requests per minute). Base URL: https://openrouter.ai/api/v1 - [Groq](https://llmperks.com/groq): Free plan with per-model limits on requests and tokens per minute and per day; very high generation speed on LPU hardware. Base URL: https://api.groq.com/openai/v1 - [Google AI Studio](https://llmperks.com/gemini): The Gemini API free tier covers Flash and Flash-Lite models with rate limits; free-tier content may be used to improve Google products. Base URL: https://generativelanguage.googleapis.com/v1beta/openai/ - [Mistral AI](https://llmperks.com/mistral): The Free plan includes monthly API credits for Mistral models, with low default rate limits. Base URL: https://api.mistral.ai/v1 - [Together AI](https://llmperks.com/together): Only selected models are priced at zero; there is no general free trial. Base URL: https://api.together.xyz/v1 - [Cohere](https://llmperks.com/cohere): Trial keys allow 1,000 API calls per month for evaluation, not for production. Base URL: https://api.cohere.ai/compatibility/v1 - [Cerebras](https://llmperks.com/cerebras): Free trial with daily token limits; a verified payment method is required to activate the API. Base URL: https://api.cerebras.ai/v1 - [Cloudflare Workers AI](https://llmperks.com/cloudflare): 10,000 Neurons per day are free on both Workers Free and Paid plans. Base URL: https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1 - [Hugging Face](https://llmperks.com/huggingface): Free accounts get a small monthly credit for Inference Providers; PRO accounts get more. Base URL: https://router.huggingface.co/v1 - [OrcaRouter](https://llmperks.com/orcarouter): BYOK router without token markup; its own Fusion models and selected free routes cost nothing. Base URL: https://api.orcarouter.ai/v1 - [NVIDIA Build](https://llmperks.com/nvidia): Free serverless inference for development and prototyping with a developer account; not intended for production traffic. Base URL: https://integrate.api.nvidia.com/v1 - [TokenRouter](https://llmperks.com/tokenrouter): Free models and promotional access may be time-limited. Base URL: https://api.tokenrouter.com/v1 - [Vercel AI Gateway](https://llmperks.com/vercel): A payment card must be on file even for zero-price models. Base URL: https://ai-gateway.vercel.sh/v1 - [Nous Portal](https://llmperks.com/nous): The $0 Portal plan includes :free models with standard rate limits. Base URL: https://inference-api.nousresearch.com/v1 - [UnoRouter](https://llmperks.com/unorouter): Large catalog of :free routes without a card; shared capacity and per-model limits. Base URL: https://api.unorouter.com/v1 ## Articles - [Free LLM APIs in 2026: the complete guide](https://llmperks.com/blog/free-llm-api-guide-2026): Where to get free API access to DeepSeek, Qwen, GLM, Kimi, Gemini, and other LLMs in 2026: 15 providers, their limits, card requirements, and how to pick a model without hitting quotas. - [How to use free LLMs in Cursor, Cline, OpenCode, and VS Code](https://llmperks.com/blog/free-models-in-cursor-cline-opencode): Step-by-step setup of free LLMs in Cline, Kilo Code, OpenCode, Continue, Aider, Zed, and Cursor: base URLs, keys, picking a coding model, and working around limits. - [Best free coding models: September 2026 ranking](https://llmperks.com/blog/best-free-coding-models-2026): Which free-API models write the best code in September 2026: GLM 5.3, Kimi K3, Qwen 3.8, DeepSeek V4, and more — benchmarks, context, and providers. - [OpenRouter free tier: :free model limits explained](https://llmperks.com/blog/openrouter-free-limits): How many requests OpenRouter's free models allow in 2026, what the $10 top-up changes, how :free routes work, and when to use the openrouter/free router. - [Hermes Agent on free models: an autonomous agent for $0](https://llmperks.com/blog/hermes-agent-free-models): How to run Nous Research's Hermes Agent with free OpenRouter or NVIDIA Build models or a local Ollama: setup, model choice, limits, and pitfalls. - [How to run an LLM locally: Ollama, LM Studio, Jan, and llama.cpp](https://llmperks.com/blog/run-llm-locally-ollama-lm-studio): Comparing free apps for local LLMs in 2026 — Ollama, LM Studio, Jan, llama.cpp — which models fit your hardware and how to connect them to your IDE. - [Free AI credits in 2026: for developers, startups, and students](https://llmperks.com/blog/free-ai-credits-2026): How to get free AI API and cloud credits in 2026: Gemini API, OpenRouter, AWS Activate, Google for Startups, Microsoft for Startups, and student programs. - [NVIDIA Nemotron 3 on OpenRouter: free API guide](https://llmperks.com/blog/nvidia-nemotron-3-openrouter): A practical guide to NVIDIA Nemotron 3 Nano, Super, Omni, and Ultra: architecture, context windows, free OpenRouter endpoints, privacy caveats, and API examples. ## Languages - English: https://llmperks.com - Deutsch: https://llmperks.com/de - Français: https://llmperks.com/fr - Español: https://llmperks.com/es - Русский: https://free-llms.ru