Run This Ai
EN DE

DeepEval

The LLM Evaluation Framework

★ 16,501 GitHub Apache-2.0 llm-evaluationtestingragllmopspython AI Agents

Overview

DeepEval is an open-source LLM evaluation framework that makes testing and evaluating LLM outputs simple and reliable. With 16.5k+ stars on GitHub, it provides a comprehensive suite of over 15 metrics for evaluating LLM responses, including G-Eval, hallucination detection, bias assessment, toxicity scoring, contextual recall, faithfulness, and more. Built for CI/CD integration from day one, DeepEval allows developers to write unit tests for their LLM applications just like they would for traditional software. It integrates seamlessly with popular frameworks like LangChain, LlamaIndex, and Guardrails, and supports advanced features like dataset management, real-time monitoring via Confident AI, and custom metric creation. Whether you are building RAG pipelines, chatbots, or agentic systems, DeepEval gives you the confidence that your LLM outputs are accurate, safe, and reliable.

Requirements

Min vCPU
1
Min RAM
512 MB
Min Disk
10 GB
Rec vCPU
2
Rec RAM
2048 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
View plan

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Affiliate disclosure

Related tools

Guides & articles