Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
★ 45,643 GitHub
MPL-2.0 text-to-speechttsspeech-synthesisvoice-cloningdeep-learningpytorchaudio Bild & Video
Überblick
Coqui TTS (ᐸ🐸ᐳ💬) is a powerful open-source deep learning toolkit for Text-to-Speech synthesis. With 45k+ GitHub stars, it supports multiple models including Tacotron, Glow-TTS, and VITS, multi-speaker TTS, voice cloning, and fine-tuning. It provides a production-ready Docker image for easy self-hosting and a comprehensive API for inference. The toolkit is battle-tested in both research and production environments.
Anforderungen
Min vCPU
2
Min RAM
2048 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB
Empfohlener VPS
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate-Hinweis
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
coqui-tts:
image: ghcr.io/coqui-ai/tts-cpu:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/coqui-tts:/data
Verwandte Tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
Bark
Text-prompted generative audio model that produces speech, music, and sound effects from natural language