ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
★ 39,525 GitHub
AGPL-3.0 text-to-speechttsvoice-synthesisconversational-aichineseenglish Image & Video
Overview
ChatTTS is a powerful open-source text-to-speech (TTS) model designed specifically for conversational scenarios. With nearly 40k GitHub stars, it generates natural, expressive speech that captures the nuances of daily dialogue including laughter, pauses, and conversational fillers. Unlike traditional TTS systems that produce robotic monotone speech, ChatTTS leverages a deep generative model trained on massive conversational data to deliver human-quality voice synthesis. It supports English and Chinese, offers voice cloning capabilities, and includes a WebUI and OpenAI-compatible API server for easy deployment. The model runs efficiently on consumer GPUs with as little as 4GB VRAM.
Requirements
Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB
Recommended VPS
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate disclosure
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
chattts:
image: tinyserve/chattts:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/chattts:/data
Related tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
Bark
Text-prompted generative audio model that produces speech, music, and sound effects from natural language