Fish Speech System Requirements (CPU, RAM, Disk)
System requirements for Fish Speech: CPU, RAM, and disk space.
Fish Speech is a cutting-edge open-source text-to-speech (TTS) system developed by Fish Audio. With over 31k GitHub stars, it delivers SOTA multilingual speech synthesis supporting 80+ languages. The S2 Pro model (4B parameters) features a Dual-Autoregressive architecture with RL alignment for exceptionally natural, emotionally rich speech. It supports sub-word level prosody control using natural language tags like [whisper], [excited], [angry], and natively handles multi-speaker and multi-turn conversation generation. Deployable via Docker with CPU and CUDA variants, it includes WebUI, server, and command-line interfaces for flexible integration.
Requirements · Fish Speech
See the full tool page: Fish Speech →