Colibri
Führen Sie Frontier-MoE-Modelle (744 Mrd. bis 2,8 Bio. Parameter) auf Ihrer vorhandenen Hardware aus. Reines C, null Abhängigkeiten, Experten werden von der Festplatte gestreamt.
Überblick
Anforderungen
Verwandte Tools
NextChat
Cross-platform ChatGPT web UI with multi-model support for Ollama, Claude, Gemini, and more
Lobe Chat
Extensible, open-source ChatGPT alternative with plugins, knowledge base, and multi-LLM support
Open WebUI
User-friendly WebUI for LLMs (Ollama, OpenAI API)
text-generation-webui
Run local LLMs with a powerful web interface — text, vision, tool-calling, and OpenAI-compatible API
Streamlit
Build and share data apps in pure Python — fast
Gradio
Build and share delightful machine learning apps in Python
Anleitungen & Artikel
How to Deploy Colibri: Chat with a 744B MoE Model on Your Own Machine
Step-by-step tutorial: build Colibri from source, chat with a 744B GLM-5.2 model from your terminal, expose it via an OpenAI-compatible API, and deploy it with Docker Compose.
Colibri: Run Frontier 2.8T MoE Models on Hardware You Already Own
Colibri is a pure-C inference engine that streams frontier MoE models — GLM-5.2 (744B), Kimi K3 (2.8T) and more — from disk onto the hardware you already own. Here's how it works and why the memory hierarchy changes everything.