Run This Ai
EN DE

CogVLM

Open-source vision-language foundation model for multimodal understanding and visual question answering

★ 6,740 GitHub Apache-2.0 vision-languagemultimodalvqaimage-captioningdeep-learningpytorchtsinghua Image & Video

Overview

CogVLM is a state-of-the-art open visual language model developed by Tsinghua University's THUDM lab. At 6.7k+ GitHub stars, it represents a significant advancement in multimodal AI. Unlike traditional approaches that use a separate visual encoder with a frozen language model, CogVLM employs a deep fusion architecture where visual and language features interact at every transformer layer. This enables superior performance on tasks like visual question answering, image captioning, and visual grounding. The model supports both image and text inputs, generating text responses. CogVLM is fully open-source under Apache-2.0, built with PyTorch, and has been adopted by researchers and developers worldwide for building multimodal AI applications.

Requirements

Min vCPU
4
Min RAM
8192 MB
Min Disk
10 GB
Rec vCPU
8
Rec RAM
16384 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Related tools

Guides & articles