Speech & Audio
DIA
Ultra-realistic dialogue TTS in one pass
β
19,358
FunASR
End-to-end speech recognition toolkit with streaming ASR, VAD, punctuation, and speaker diarization
β
19,532
Supertonic
Lightning-fast, on-device, multilingual TTS running natively via ONNX.
β
13,564
Abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
β
5,461
ebook2audiobook
Generate audiobooks from e-books with voice cloning and 1158+ languages.
β
19,620
NeMo Speech
NVIDIA scalable generative AI framework for speech β ASR, TTS, and speech language models in one toolkit.
β
17,871
ESPnet
The end-to-end speech processing toolkit for ASR, TTS, speech translation, and speaker diarization.
β
9,912