Lemonade
Run optimized LLMs locally on your own GPU or NPU — a fast, OpenAI-compatible server for private AI apps.
Overview
Requirements
Recommended VPS
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate disclosure
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
lemonade:
image: ghcr.io/lemonade-sdk/lemonade:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/lemonade:/data
Related tools
NextChat
Cross-platform ChatGPT web UI with multi-model support for Ollama, Claude, Gemini, and more
Lobe Chat
Extensible, open-source ChatGPT alternative with plugins, knowledge base, and multi-LLM support
Open WebUI
User-friendly WebUI for LLMs (Ollama, OpenAI API)
text-generation-webui
Run local LLMs with a powerful web interface — text, vision, tool-calling, and OpenAI-compatible API
Streamlit
Build and share data apps in pure Python — fast
Gradio
Build and share delightful machine learning apps in Python
Guides & articles
Getting Started with Lemonade: Local LLMs in Minutes
Step-by-step tutorial: deploy Lemonade with Docker, download an optimized local LLM, and connect your apps via the OpenAI-compatible API — fully offline.
Lemonade: Run Local AI Apps on Your Own GPU
Lemonade is an open-source local AI server that runs optimized LLMs on your own GPU or NPU. Discover how it delivers private, OpenAI-compatible inference with 5.4K+ GitHub stars.