CogVLM
Open-source vision-language foundation model for multimodal understanding and visual question answering
Overview
Requirements
Recommended VPS
Hostinger · KVM 8
8 vCPU · 32256 MB · 400 GB
Affiliate disclosure
Related tools
Stable Diffusion WebUI
Generative image models in your browser
ComfyUI
The most powerful and modular diffusion model GUI with a graph/nodes interface for Stable Diffusion
whisper.cpp
High-performance C++ port of OpenAI Whisper for fast local speech recognition
Fooocus
AI image generator focusing on prompts and generating — a Midjourney-like experience offline
Coqui TTS
Open-source deep learning toolkit for text-to-speech, battle-tested in research and production
ChatTTS
High-quality conversational text-to-speech model optimized for natural daily dialogue
Guides & articles
How to Run CogVLM Locally: A Step-by-Step Tutorial
Step-by-step guide to setting up and running CogVLM locally. From cloning the repository to running inference with the web demo and Python API.
CogVLM: Tsinghua's Open-Source Vision-Language Model with Deep Fusion Architecture
Discover CogVLM, the open-source vision-language model from Tsinghua University's THUDM lab. Learn about its deep fusion architecture, key features, and use cases for multimodal AI.