Bark
Text-prompted generative audio model that produces speech, music, and sound effects from natural language
★ 39,182 GitHub
MIT text-to-speechaudio-generationmusic-generationsound-effectsmultilingualsuno Bild & Video
Überblick
Bark is a transformer-based text-prompted generative audio model developed by Suno AI. With nearly 40,000 GitHub stars, it can generate highly realistic and multilingual speech, music, background noises, and sound effects from simple text prompts. Unlike traditional TTS systems, Bark understands context and can produce non-verbal sounds like laughs, sighs, and throat clears. It supports 13+ languages and can generate music with vocals, making it one of the most versatile open-source audio generation models available. Bark runs on consumer GPUs with as little as 4GB VRAM and offers both Python API and Docker deployment options.
Anforderungen
Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB
Empfohlener VPS
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate-Hinweis
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
bark:
image: barksim/bark:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/bark:/data
Verwandte Tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue