Run This Ai
EN DE

LLaVA

Large Language and Vision Assistant — multimodal AI with GPT-4 level capabilities

★ 25,000 GitHub Apache-2.0 vision-languagemultimodalllmimage-understandingai-assistantneural-networkdeep-learning Image & Video

Overview

LLaVA (Large Language and Vision Assistant) is a groundbreaking open-source multimodal AI model that combines vision and language understanding with GPT-4 level capabilities. Developed by researchers from UW-Madison and Microsoft, LLaVA achieves state-of-the-art performance on 11 benchmarks through visual instruction tuning. It can understand images, answer questions about visual content, and engage in natural conversations about what it sees. With 25k+ GitHub stars and Apache-2.0 licensing, LLaVA supports LLama-3, Qwen-1.5, and various model sizes from 7B to 110B parameters. The model can be fine-tuned with LoRA on consumer GPUs and deployed for both research and production use cases.

Requirements

Min vCPU
2
Min RAM
8192 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
16384 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Related tools

Guides & articles