MetaVoice
Foundational model for human-like, expressive text-to-speech (TTS) with emotional speech rhythm and zero-shot voice cloning
Overview
Requirements
Recommended VPS
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate disclosure
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
metavoice:
image: saladtechnologies/metavoice-api:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/metavoice:/data
Related tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
Guides & articles
How to Run MetaVoice-1B with Docker: A Step-by-Step Guide
Learn how to deploy MetaVoice-1B TTS model with Docker in minutes. Step-by-step guide covering installation, API usage, voice cloning, Web UI, and troubleshooting.
MetaVoice-1B: Human-like Expressive TTS with Emotional Voice Cloning
MetaVoice-1B is a 1.2B parameter open-source TTS model with emotional speech, zero-shot voice cloning in 30 seconds, and cross-lingual cloning with 1 minute of data — all under Apache 2.0 license.