v11.9.065892.7 MB
MIT
strict
core24
Local AI server with OpenAI-compatible API
Lemonade is a local AI server that provides cloud-API-equivalent capabilities
(chat, coding, speech, image generation) running 100% locally on your own
hardware. It exposes OpenAI, Anthropic, and Ollama-compatible APIs on
localhost:13305, connectable to hundreds of apps.
Features:
- OpenAI, Anthropic, and Ollama-compatible APIs (chat completions, embeddings, etc.)
- Multi-modal inference: LLMs, speech-to-text, text-to-speech, audio generation,
image generation, and 3D generation
- Automatic backend selection based on available hardware — CUDA, ROCm, Vulkan,
NPU, and CPU backends download on demand
- 77+ built-in models (GGUF, FLM, ONNX) with
for model management
- Cloud offload: route inference to OpenAI-compatible providers alongside
local models
- MCP (Model Context Protocol) client and server support
- Built-in Prometheus metrics endpoint
- Runs as a background service with auto-restart on failure
Supported Hardware:
- AMD GPUs: RDNA2/3/4 (RX 6000/7000/9000), Strix Point/Halo APUs, Instinct MI100/MI200/MI300X/MI350X via ROCm
- AMD XDNA2 NPUs (Ryzen AI) for LLM and speech inference
- NVIDIA GPUs: Turing (sm75) and newer via CUDA (Ampere, Ada Lovelace, Hopper, Blackwell)
- Cross-vendor GPUs via Vulkan (AMD, NVIDIA, Intel, Qualcomm Adreno on ARM64)
- CPU fallback for systems without GPU acceleration
- Both amd64 (x8664) and arm64 (aarch64) architectures
Quick Start:
The server starts automatically after installation. Access the API at: http://localhost:13305/api/v1
Documentation: https://lemonade-server.ai/
(chat, coding, speech, image generation) running 100% locally on your own
hardware. It exposes OpenAI, Anthropic, and Ollama-compatible APIs on
localhost:13305, connectable to hundreds of apps.
Features:
- OpenAI, Anthropic, and Ollama-compatible APIs (chat completions, embeddings, etc.)
- Multi-modal inference: LLMs, speech-to-text, text-to-speech, audio generation,
image generation, and 3D generation
- Automatic backend selection based on available hardware — CUDA, ROCm, Vulkan,
NPU, and CPU backends download on demand
- 77+ built-in models (GGUF, FLM, ONNX) with
lemonade-server pull / lemonade-server listfor model management
- Cloud offload: route inference to OpenAI-compatible providers alongside
local models
- MCP (Model Context Protocol) client and server support
- Built-in Prometheus metrics endpoint
- Runs as a background service with auto-restart on failure
Supported Hardware:
- AMD GPUs: RDNA2/3/4 (RX 6000/7000/9000), Strix Point/Halo APUs, Instinct MI100/MI200/MI300X/MI350X via ROCm
- AMD XDNA2 NPUs (Ryzen AI) for LLM and speech inference
- NVIDIA GPUs: Turing (sm75) and newer via CUDA (Ampere, Ada Lovelace, Hopper, Blackwell)
- Cross-vendor GPUs via Vulkan (AMD, NVIDIA, Intel, Qualcomm Adreno on ARM64)
- CPU fallback for systems without GPU acceleration
- Both amd64 (x8664) and arm64 (aarch64) architectures
Quick Start:
The server starts automatically after installation. Access the API at: http://localhost:13305/api/v1
Documentation: https://lemonade-server.ai/
Update History
v11.8.1 (602) → v11.9.0 (658)3 Sept 2026, 01:15 UTC
v11.8.0 (590) → v11.8.1 (602)28 Aug 2026, 14:30 UTC
v11.7.0 (519) → v11.8.0 (590)26 Aug 2026, 22:15 UTC
v11.6.0 (460) → v11.7.0 (519)20 Aug 2026, 16:45 UTC
v11.5.2 (378) → v11.6.0 (460)17 Aug 2026, 19:15 UTC
v11.5.1 (360) → v11.5.2 (378)6 Aug 2026, 00:30 UTC
v11.5.1 356 → 36030 Jul 2026, 14:00 UTC
v11.5.0 (351) → v11.5.1 (356)30 Jul 2026, 01:15 UTC
v11.5.0 337 → 35128 Jul 2026, 23:15 UTC
v11.0.0 (322) → v11.5.0 (337)22 Jul 2026, 21:45 UTC
v10.10.0 (310) → v11.0.0 (322)15 Jul 2026, 22:15 UTC
v10.9.0 (293) → v10.10.0 (310)8 Jul 2026, 19:00 UTC
v10.8.1 (274) → v10.9.0 (293)2 Jul 2026, 18:45 UTC
v10.8.0 (257) → v10.8.1 (274)25 Jun 2026, 15:00 UTC
v10.7.0 (232) → v10.8.0 (257)18 Jun 2026, 00:30 UTC
v10.6.0 (193) → v10.7.0 (232)11 Jun 2026, 21:45 UTC
v10.5.1 (190) → v10.6.0 (193)3 Jun 2026, 13:45 UTC
v10.5.0 (185) → v10.5.1 (190)20 May 2026, 15:30 UTC
v10.4.0 (183) → v10.5.0 (185)18 May 2026, 03:00 UTC
v10.4.0 177 → 18317 May 2026, 01:30 UTC
14 Jan 2026, 20:51 UTC
2 Sept 2026, 18:59 UTC
15 Jan 2026, 04:37 UTC