Run This Ai
EN DE

DeepEval

The LLM Evaluation Framework

★ 16,501 GitHub Apache-2.0 llm-evaluationtestingragllmopspython KI-Agenten

Überblick

DeepEval is an open-source LLM evaluation framework that makes testing and evaluating LLM outputs simple and reliable. With 16.5k+ stars on GitHub, it provides a comprehensive suite of over 15 metrics for evaluating LLM responses, including G-Eval, hallucination detection, bias assessment, toxicity scoring, contextual recall, faithfulness, and more. Built for CI/CD integration from day one, DeepEval allows developers to write unit tests for their LLM applications just like they would for traditional software. It integrates seamlessly with popular frameworks like LangChain, LlamaIndex, and Guardrails, and supports advanced features like dataset management, real-time monitoring via Confident AI, and custom metric creation. Whether you are building RAG pipelines, chatbots, or agentic systems, DeepEval gives you the confidence that your LLM outputs are accurate, safe, and reliable.

Anforderungen

Min vCPU
1
Min RAM
512 MB
Min Disk
10 GB
Rec vCPU
2
Rec RAM
2048 MB
Rec Disk
20 GB

Empfohlener VPS

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
Zum Anbieter

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
Zum Anbieter

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
Zum Anbieter

Affiliate-Hinweis

Verwandte Tools

Anleitungen & Artikel