Run This Ai
EN DE

DS4

Local inference engine for DeepSeek 4 Flash and PRO — run on Metal, CUDA and ROCm

★ 21,517 GitHub MIT llminferencedeepseekmetalcudarocmlocal-aioffline LLM & Chat

Overview

DS4 lets you run DeepSeek 4 Flash and PRO models fully offline on your own hardware. Supports Apple Metal, NVIDIA CUDA and AMD ROCm with a simple setup, automatic model downloads and an OpenAI-compatible local API.

Requirements

Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  ds4:
    image: arraying/ds4:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/ds4:/data

Best VPS for DS4 →

Related tools

Guides & articles