Run This Ai
EN DE

Aphrodite Engine

High-performance LLM inference engine for roleplay and chatbots

★ 1,771 GitHub AGPL-3.0 llminferencevllmroleplayquantizationself-hostedapi-server LLM & Chat

Überblick

Aphrodite Engine is a high-performance inference engine built on vLLM's Paged Attention technology, optimized for serving HuggingFace-compatible models at scale. It powers PygmalionAI's chat platforms and API infrastructure. Aphrodite supports an extensive range of quantization methods including AWQ, GPTQ, GGUF, ExLlamaV3, Bitsandbytes, and more. It features continuous batching, efficient K/V cache management, distributed inference, speculative decoding (EAGLE, DFlash, MTP), multi-LoRA support, multimodal support, and modern samplers like DRY, XTC, and Mirostat.

Anforderungen

Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
8
Rec RAM
16384 MB
Rec Disk
20 GB

Empfohlener VPS

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
Zum Anbieter

Affiliate-Hinweis

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  aphrodite-engine:
    image: ghcr.io/pygmalionai/aphrodite-engine:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/aphrodite-engine:/data

Bester VPS für Aphrodite Engine →

Verwandte Tools

Anleitungen & Artikel