Run This Ai
EN DE

Image & Video

Generative image, audio and video models.

Stable Diffusion WebUI

Generative image models in your browser

β˜… 140,000

CogVLM

Open-source vision-language foundation model for multimodal understanding and visual question answering

β˜… 6,740

MeloTTS

High-quality multilingual text-to-speech by MyShell.ai β€” supports English, Spanish, French, Chinese, Japanese, and Korean

β˜… 7,518

SwarmUI

Modular Stable Diffusion Web UI with ComfyUI backend β€” high performance, extensible

β˜… 4,279

Open-Sora

Democratizing efficient video production with open-source AI

β˜… 29,156

Kohya SS

GUI for Stable Diffusion training β€” LoRA, Dreambooth, and fine-tuning made easy

β˜… 12,421

CosyVoice

Alibaba's multilingual voice generation and cloning with natural emotion, tone, and accent control

β˜… 21,898

Tortoise TTS

High-quality multi-voice text-to-speech with emphasis on natural prosody, emotion, and voice variety

β˜… 14,865

F5-TTS

High-quality non-autoregressive text-to-speech with flow matching

β˜… 14,848

SenseVoice

Multilingual speech understanding β€” ASR, emotion recognition & audio event detection, 50+ languages

β˜… 8,734

moondream

Tiny open-source vision-language model that runs on CPU, in ~2GB RAM β€” image captioning, VQA, OCR

β˜… 9,813

Demucs

Meta's AI music source separation β€” split any song into vocals, drums, bass, and other stems

β˜… 10,263

Pixelle Video

AI-powered fully automated short video engine β€” generate viral shorts from scripts or topics with AI voiceovers, captions, and effects

β˜… 24,217

Toonflow

Open-source AI short drama creator β€” turn novels and scripts into animated short dramas with AI scriptwriting, storyboarding, character and video generation

β˜… 11,316

MetaVoice

Foundational model for human-like, expressive text-to-speech (TTS) with emotional speech rhythm and zero-shot voice cloning

β˜… 4,202

Scriberr

Self-hosted AI audio transcription with speaker detection, AI chat, and pristine privacy. Fully offline, no data ever leaves your server.

β˜… 2,840

OpenPencil

Open-source AI-native vector design tool with concurrent Agent Teams. Turn prompts into UI directly on the live canvas. Design-as-Code β€” a modern, self-hostable alternative to Pencil.

β˜… 4,634

Sana

Efficient high-resolution image synthesis with a linear diffusion transformer β€” 4K images up to 100x faster.

β˜… 8,792

FastVideo

A unified inference & post-training framework for accelerated video generation

β˜… 3,985

LightX2V

Lightweight Image, Video & Action Generation Inference Framework

β˜… 2,714

GPT Image Playground

Self-hosted image generation & editing playground powered by OpenAI gpt-image-2 API β€” text-to-image, reference editing & mask inpainting.

β˜… 3,523

LLaVA

Large Language and Vision Assistant β€” multimodal AI with GPT-4 level capabilities

β˜… 25,000

SnapOtter

SnapOtter

β˜… 1,797

ComfyUI

The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion

β˜… 118,476

InvokeAI

Professional creative AI tools for visual media β€” generate and edit images with an industry-leading UI

β˜… 23,000

Fish Speech

State-of-the-art open-source multilingual text-to-speech and voice cloning system

β˜… 31,004

Spleeter

Deezer source separation library β€” isolate vocals, drums, bass & more with AI

β˜… 28,255

Fooocus

AI image generator focusing on prompts and generating β€” a Midjourney-like experience offline

β˜… 50,553

OpenVoice

Instant voice cloning by MIT and MyShell - Audio foundation model

β˜… 36,798

Fooocus-API

REST API for Fooocus – Generate stunning AI images via API endpoints

β˜… 668

Buzz

Transcribe and translate audio offline on your PC. Powered by OpenAI Whisper.

β˜… 19,862

GFPGAN

AI-powered face restoration β€” enhance blurry old faces with remarkable quality

β˜… 37,481

Coqui TTS

Open-source deep learning toolkit for text-to-speech, battle-tested in research and production

β˜… 45,643

whisper.cpp

High-performance C++ port of OpenAI Whisper for fast local speech recognition

β˜… 51,133

ChatTTS

High-quality conversational text-to-speech model optimized for natural daily dialogue

β˜… 39,525

Bark

Text-prompted generative audio model that produces speech, music, and sound effects from natural language

β˜… 39,182

RVC

Retrieval-based voice conversion WebUI β€” train AI voice models with just 10 minutes of audio

β˜… 36,196

WhisperX

Automatic speech recognition with word-level timestamps and speaker diarization

β˜… 22,785

Real-ESRGAN

Image super-resolution upscaler with AI β€” enhance and enlarge images with remarkable quality

β˜… 35,946

Faster-Whisper

Fast Whisper transcription with CTranslate2 β€” 4x faster than OpenAI Whisper with lower memory usage

β˜… 23,921

AudioCraft

Meta's open-source AI library for audio generation β€” MusicGen for text-to-music and AudioGen for text-to-sound

β˜… 23,426