Run This Ai
EN DE

NeMo Speech

NVIDIA scalable generative AI framework for speech — ASR, TTS, and speech language models in one toolkit.

★ 17,871 GitHub Apache-2.0 speechasrttsnvidianemoaudio Speech & Audio

Overview

NeMo Speech (NVIDIA NeMo) is a scalable generative AI framework built for researchers and developers working on speech AI. It provides production-ready building blocks for automatic speech recognition (ASR), text-to-speech (TTS), and speech language models, including tools like Parakeet and Canary for high-accuracy transcription and translation, and FastConformer-based models for real-time inference. Apache-2.0 licensed with over 17K GitHub stars.

Requirements

Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  nemo-speech:
    image: nvcr.io/nvidia/nemo:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/nemo-speech:/data

Best VPS for NeMo Speech →

Related tools

Guides & articles