Run This Ai
EN DE

Lemonade

Run optimized LLMs locally on your own GPU or NPU — a fast, OpenAI-compatible server for private AI apps.

★ 5,385 GitHub Apache-2.0 local-aillmgpunpuonnxopenai-apiprivacy LLM & Chat

Overview

Lemonade helps you discover and run local AI apps by serving optimized LLMs directly from your own GPU or NPU. It is a lightweight, OpenAI-compatible local server that keeps your data private while delivering fast inference on consumer hardware. Built on ONNX Runtime with Vulkan, ROCm, and NPU acceleration, it supports models like Llama, Mistral, and Qwen without cloud dependencies.

Requirements

Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  lemonade:
    image: ghcr.io/lemonade-sdk/lemonade:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/lemonade:/data

Best VPS for Lemonade →

Related tools

Guides & articles