LLaVA
Large Language and Vision Assistant — multimodal AI with GPT-4 level capabilities
Overview
Requirements
Recommended VPS
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate disclosure
Related tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
Guides & articles
Getting Started with LLaVA: Run Multimodal AI on Your Own Machine
Learn to run LLaVA locally: llama.cpp, official repo, or Gradio UI. Supports consumer GPUs with 4-bit quantization.
LLaVA: The Open-Source Vision-Language Model Powering Multimodal AI
Discover LLaVA, the open-source vision-language model bringing GPT-4 level multimodal understanding to everyone. 25k+ stars, Apache-2.0.