Lemonade
Local AI server for text, images and speech with an OpenAI-compatible API, optimized for AMD Ryzen AI NPUs and GPUs.
- License
- Apache-2.0
- Platforms
- Windows, macOS, Linux, Docker, CLI
- Own API keys
- Yes
- Checked
- 2026-09-24
Lemonade works with OpenAI-compatible APIs. Grab a free key and a model from the catalog: for coding · for agents · providers and base URLs
Lemonade alternatives
Popular runtime to download and run open LLMs locally via CLI, desktop app and a REST/OpenAI-compatible API.
Desktop app to discover, download and run local LLMs (llama.cpp, MLX) with a chat UI and an OpenAI-compatible local server.
Offline ChatGPT alternative: runs local LLMs, connects to cloud APIs, supports MCP and serves an OpenAI-compatible API on localhost.
Reference C/C++ engine for GGUF models; llama-server provides an OpenAI-compatible API and web UI on CPU, CUDA, Metal, Vulkan and ROCm.
High-throughput LLM inference and serving engine with an OpenAI-compatible server for NVIDIA, AMD, Intel, TPU and CPU.
Self-hosted drop-in OpenAI/Anthropic API replacement for LLMs, images, audio, video and embeddings on CPU or GPU.
FAQ
Is Lemonade free?
Completely free and open source, with zero telemetry.
Does Lemonade work with free models?
Yes. It supports your own keys and OpenAI-compatible APIs, so you can plug in free models from OpenRouter, Groq, NVIDIA Build, or a local Ollama.
What license does Lemonade use?
Apache-2.0. The source code is open.