Speech & Audio
DIA
Ultra-realistic dialogue TTS in one pass
FunASR
End-to-end speech recognition toolkit with streaming ASR, VAD, punctuation, and speaker diarization
Supertonic
Lightning-fast, on-device, multilingual TTS running natively via ONNX.
Abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
ebook2audiobook
Generate audiobooks from e-books with voice cloning and 1158+ languages.
NeMo Speech
NVIDIA scalable generative AI framework for speech β ASR, TTS, and speech language models in one toolkit.
ESPnet
The end-to-end speech processing toolkit for ASR, TTS, speech translation, and speaker diarization.
Piper
Fast, local neural text-to-speech that runs fully offline
Moonshine Voice
Real-time on-device voice AI: speech-to-text, intent recognition and text-to-speech for building fast, private voice agents and interfaces.
Anarlog
Local-first AI meeting notetaker β transcribe on-device, bring your own LLM
Ultravox
Open-source multimodal LLM for real-time voice AI β text in, speech out with ultra-low latency.
Vexa
Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom
WhisperLiveKit
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.