Xinference
Run open-source LLMs, embeddings, and multimodal models with one line of code
Überblick
Anforderungen
Empfohlener VPS
Hostinger · KVM 2
2 vCPU · 8192 MB · 100 GB
Hostinger · KVM 2
2 vCPU · 8192 MB · 100 GB
Hostinger · KVM 4
4 vCPU · 16384 MB · 200 GB
Affiliate-Hinweis
Docker Compose
# Generated by Run This Ai — docker-compose.yml
services:
xinference:
image: xprobe/xinference:latest
restart: unless-stopped
ports:
- 8080:8080
volumes:
- ./data/xinference:/data
Verwandte Tools
NextChat
Cross-platform ChatGPT web UI with multi-model support for Ollama, Claude, Gemini, and more
Lobe Chat
Extensible, open-source ChatGPT alternative with plugins, knowledge base, and multi-LLM support
Open WebUI
User-friendly WebUI for LLMs (Ollama, OpenAI API)
text-generation-webui
Run local LLMs with a powerful web interface — text, vision, tool-calling, and OpenAI-compatible API
Streamlit
Build and share data apps in pure Python — fast
Gradio
Build and share delightful machine learning apps in Python
Anleitungen & Artikel
Xinference vs TGI: Choosing the Right LLM Server for Self-Hosting
Detailed comparison of Xinference vs Text Generation Inference (TGI): model support, performance, features, and recommendations for choosing the right LLM server.
Getting Started with Xinference: One CLI to Rule All LLMs
Complete guide to getting started with Xinference: deploy LLMs, embeddings, and multimodal models with one CLI command. Docker setup, pip install, and LangChain integration.