Image & Video
Generative image, audio and video models.
Stable Diffusion WebUI
Generative image models in your browser
CogVLM
Open-source vision-language foundation model for multimodal understanding and visual question answering
MeloTTS
High-quality multilingual text-to-speech by MyShell.ai β supports English, Spanish, French, Chinese, Japanese, and Korean
SwarmUI
Modular Stable Diffusion Web UI with ComfyUI backend β high performance, extensible
Open-Sora
Democratizing efficient video production with open-source AI
Kohya SS
GUI for Stable Diffusion training β LoRA, Dreambooth, and fine-tuning made easy
CosyVoice
Alibaba's multilingual voice generation and cloning with natural emotion, tone, and accent control
Tortoise TTS
High-quality multi-voice text-to-speech with emphasis on natural prosody, emotion, and voice variety
F5-TTS
High-quality non-autoregressive text-to-speech with flow matching
SenseVoice
Multilingual speech understanding β ASR, emotion recognition & audio event detection, 50+ languages
moondream
Tiny open-source vision-language model that runs on CPU, in ~2GB RAM β image captioning, VQA, OCR
Demucs
Meta's AI music source separation β split any song into vocals, drums, bass, and other stems
Pixelle Video
AI-powered fully automated short video engine β generate viral shorts from scripts or topics with AI voiceovers, captions, and effects
Toonflow
Open-source AI short drama creator β turn novels and scripts into animated short dramas with AI scriptwriting, storyboarding, character and video generation
MetaVoice
Foundational model for human-like, expressive text-to-speech (TTS) with emotional speech rhythm and zero-shot voice cloning
Scriberr
Self-hosted AI audio transcription with speaker detection, AI chat, and pristine privacy. Fully offline, no data ever leaves your server.
OpenPencil
Open-source AI-native vector design tool with concurrent Agent Teams. Turn prompts into UI directly on the live canvas. Design-as-Code β a modern, self-hostable alternative to Pencil.
Sana
Efficient high-resolution image synthesis with a linear diffusion transformer β 4K images up to 100x faster.
FastVideo
A unified inference & post-training framework for accelerated video generation
LightX2V
Lightweight Image, Video & Action Generation Inference Framework
GPT Image Playground
Self-hosted image generation & editing playground powered by OpenAI gpt-image-2 API β text-to-image, reference editing & mask inpainting.
LLaVA
Large Language and Vision Assistant β multimodal AI with GPT-4 level capabilities
SnapOtter
SnapOtter
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
InvokeAI
Professional creative AI tools for visual media β generate and edit images with an industry-leading UI
Fish Speech
State-of-the-art open-source multilingual text-to-speech and voice cloning system
Spleeter
Deezer source separation library β isolate vocals, drums, bass & more with AI
Fooocus
AI image generator focusing on prompts and generating β a Midjourney-like experience offline
OpenVoice
Instant voice cloning by MIT and MyShell - Audio foundation model
Fooocus-API
REST API for Fooocus β Generate stunning AI images via API endpoints
Buzz
Transcribe and translate audio offline on your PC. Powered by OpenAI Whisper.
GFPGAN
AI-powered face restoration β enhance blurry old faces with remarkable quality
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
Bark
Text-prompted generative audio model that produces speech, music, and sound effects from natural language
RVC
Retrieval-based voice conversion WebUI β train AI voice models with just 10 minutes of audio
WhisperX
Automatic speech recognition with word-level timestamps and speaker diarization
Real-ESRGAN
Image super-resolution upscaler with AI β enhance and enlarge images with remarkable quality
Faster-Whisper
Fast Whisper transcription with CTranslate2 β 4x faster than OpenAI Whisper with lower memory usage
AudioCraft
Meta's open-source AI library for audio generation β MusicGen for text-to-music and AudioGen for text-to-sound