LLM & Chat
Selbst-gehostete LLM-Server und Chat-Oberflächen.
Open WebUI
User-friendly WebUI for LLMs (Ollama, OpenAI API)
AxonHub
Open-Source-KI-Gateway – nutzen Sie jedes SDK, um über 100 LLMs mit integriertem Failover, Lastverteilung und Kostenkontrolle aufzurufen.
EvalScope
Ein schlankes und anpassbares Framework für die effiziente Evaluierung und das Performance-Benchmarking großer Modelle (LLMs, VLMs, KI-Agenten).
Rig
Entwickle modulare und skalierbare LLM-Anwendungen in Rust
H2O LLM Studio
Ein No-Code-GUI-Framework für das Fine-Tuning von Large Language Models mit intuitivem Dashboard und Experiment-Tracking
WindsurfAPI
Selbstgehostetes KI-Modell-Gateway für über 100 Modelle
DEEIX Chat
Enterprise-KI-Arbeitsbereich mit Modell-Routing, multimodalem Chat, Dateien, Tools und Abrechnungsmanagement
DeepClaude
Hochleistungs-LLM-Inferenz-API, die das Schlussfolgerungsvermögen von DeepSeek R1 mit der kreativen Kraft von Claude kombiniert
CoAI
Multi-Tenant-KI-Plattform der nächsten Generation mit LLM-Gateway auf Unternehmensniveau, integrierter Abrechnung und Unterstützung für über 200 KI-Modelle
AI as Workspace
Eleganter KI-Chat-Client mit Multi-Workspace-Unterstützung, Plugin-System, MCP-Integration und Echtzeit-Cloud-Synchronisierung
Evidently
Open-Source-Framework für ML- und LLM-Observability. Evaluieren, testen und überwachen Sie jedes KI-gestützte System oder jede Daten-Pipeline mit über 100 Metriken.
MLflow
Open-Source-KI-Engineering-Plattform für Agenten, LLMs und ML-Modelle. Debuggen, evaluieren, überwachen und optimieren Sie KI-Anwendungen in Produktionsqualität.
OmniRoute
Kostenloses KI-Gateway: ein Endpunkt, uber 231 Anbieter (uber 50 kostenlos), verbinde Claude Code, Codex, Cursor, Cline & Copilot mit KOSTENLOSEN Claude/GPT/Gemini-Modellen. RTK+Caveman-Komprimierung spart 15-95 % an Tokens.
Manifest
Der Open-Source-LLM-Router, der Ihre KI-Agenten in Sekundenschnelle mit jedem Anbieter verbindet. Abonnements, Pay-per-Token, lokale Modelle und benutzerdefinierte Anbieter.
FastChat
Open-source platform for training, serving, and evaluating large language models — from the makers of Vicuna and Chatbot Arena
FastAgency
The fastest way to bring multi-agent workflows to production
SillyTavern
LLM Frontend for Power Users - roleplay, chat, and AI interactions
Lemonade
Führe optimierte LLMs lokal auf deiner eigenen GPU oder NPU aus – ein schneller, OpenAI-kompatibler Server für private KI-Anwendungen.
OpenMed
Local-first Gesundheits-KI für klinische NER und HIPAA-konforme PII-Deidentifizierung, die zu 100 % auf Ihrer eigenen Hardware läuft.
Giskard
Open-Source-Evaluierungs- und Testbibliothek für LLM-Agenten
RLLM
Agentisches RL auf jedem Harness, mit jedem Backend, auf jedem Benchmark.
Soup
Feinabstimmung von LLMs über eine einzige YAML-Datei. Layer-Streaming trainiert ein 8B-Modell auf einer 4-GB-Laptop-GPU.
ODS
Verwandle deinen PC, Mac oder Linux-Rechner in einen vollwertigen KI-Server – LLM-Inferenz, Chat-UI, Sprache, Agenten, RAG und Workflows.
vLLM Omni
Ein auf vLLM basierendes Framework für die effiziente Modellinferenz mit Omni-Modalitäts-Modellen.
DS4
Lokale Inferenz-Engine für DeepSeek 4 Flash und PRO – ausführbar auf Metal, CUDA und ROCm
Colibri
Führen Sie Frontier-MoE-Modelle (744 Mrd. bis 2,8 Bio. Parameter) auf Ihrer vorhandenen Hardware aus. Reines C, null Abhängigkeiten, Experten werden von der Festplatte gestreamt.
Chainlit
Build Conversational AI chatbots in minutes with Python
GPUStack
GPU-Cluster verwalten, KI-Modelle mit vLLM und SGLang bereitstellen und auf Abruf SSH-zugängliche GPU-Instanzen erhalten.
vLLM Ascend
High-Throughput LLM-Serving auf Huawei Ascend NPUs. Community-gepflegtes Hardware-Plugin für vLLM.
Mesh LLM
Dezentrale KI/LLM für alle. Teile deine Rechenleistung privat oder öffentlich, um deine Agenten und Chats zu betreiben.
Bifrost
Das schnellste Enterprise-AI-Gateway – 50x schneller als LiteLLM, mit adaptivem Load Balancing, Clustermodus, Guardrails und Unterstützung für über 1000 Modelle.
BentoML
Der einfachste Weg, KI-Apps und -Modelle bereitzustellen – erstellen Sie Inference-APIs, Job-Queues, LLM-Apps und Multi-Modell-Pipelines.
SeekDB
Die KI-native Suchdatenbank. Vereint Vektor-, Volltext- und skalare Suche in einer einzigen Engine – der beste Speicher für KI-Agenten, RAG-Pipelines und agentische Workflows.
OpenMetadata
Die offene Plattform für vertrauenswürdigen Datenkontext und Geschäftssemantik für Menschen, KI-Assistenten und Agenten.
FauxPilot
Open-source alternative to Copilot using Triton Inference Server
Xinference
Run open-source LLMs, embeddings, and multimodal models with one line of code
Text Generation Inference
Hugging Face's high-performance LLM serving with Rust/Python for production
TabbyAPI
Lightweight OpenAI-compatible ExLlamaV2 API server
Aphrodite Engine
High-performance LLM inference engine for roleplay and chatbots
Llamafile
Distribute and run LLMs as single-file executables — no installation needed
LMDeploy
Efficient LLM deployment with TurboMind/PyTorch engines
SGLang
High-performance serving framework for LLMs with RadixAttention
Intel transformers
Intel extension for transformers and LLM serving
text-generation-webui
Run local LLMs with a powerful web interface — text, vision, tool-calling, and OpenAI-compatible API
aisuite
Simple, unified interface to multiple generative AI providers
Open WebUI Pipelines
Connection framework for Open WebUI — add custom filters, function routing, and pipelines to any LLM
LiteLLM
Open-source AI gateway to call 100+ LLM providers in OpenAI format — self-hosted, enterprise-ready
LocalAI
Open-source AI engine — run any model (LLM, vision, voice, image, video) on any hardware, no GPU required
Jan
Open-source ChatGPT replacement — run LLMs locally with full control and privacy
LibreChat
Self-hosted AI chat platform unifying all major LLM providers in one privacy-focused UI
Phoenix
AI observability & evaluation: LLM tracing, evaluation, and RAG troubleshooting
Streamlit
Build and share data apps in pure Python — fast
OpenVINO
Intel's open-source toolkit for optimizing and deploying AI inference across hardware platforms
Lobe Chat
Extensible, open-source ChatGPT alternative with plugins, knowledge base, and multi-LLM support
llama-cpp-python
Python bindings for llama.cpp with OpenAI-compatible server and multi-model support
Langfuse
Open source AI engineering platform for LLM observability, evals, and prompt management
Bionic GPT
On-premise replacement for ChatGPT with enterprise data confidentiality
LangServe
Deploy LangChain runnables and chains as production-ready REST APIs
DSPy
The framework for programming—not prompting—language models
NextChat
Cross-platform ChatGPT web UI with multi-model support for Ollama, Claude, Gemini, and more
Helicone
Open-source LLM observability platform for monitoring, evaluating, and improving your AI applications
HuggingChat
Open-source chat interface powering HuggingChat, built by Hugging Face
Big-AGI
The Expert's AI Workspace — multi-model reasoning, personas, and advanced agent controls
Open Assistant
Open-source chat-based assistant that understands tasks and interacts with third-party systems
Chatbot UI Lite
Extremely lightweight ChatGPT-style chat interface built with Next.js, TypeScript and Tailwind CSS
Gradio
Build and share delightful machine learning apps in Python
Chatbot UI
The open-source AI chat app for everyone — chat with 80+ AI models