Run This Ai
EN DE

EvalScope

Streamlined and customizable framework for efficient large model (LLM, VLM, AI agent) evaluation and performance benchmarking.

★ 3,183 GitHub Apache-2.0 llmevaluationbenchmarkvlmagents LLM & Chat

Overview

EvalScope is a streamlined, customizable framework from Alibaba's ModelScope team for efficient evaluation and performance benchmarking of large models — covering LLMs, VLMs, and AI agents. It provides a unified API for running benchmarks like MMLU, GSM8K, and HumanEval, supports native and OpenAI-compatible model APIs, and integrates with the ModelScope ecosystem. Evaluate models across multiple benchmarks with a single command, generate detailed reports, and compare results side by side.

Requirements

Min vCPU
2
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
4
Rec RAM
8192 MB
Rec Disk
20 GB

Recommended VPS

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
View plan

Hostinger · KVM 8

8 vCPU · 32256 MB · 400 GB

$20.00
View plan

Affiliate disclosure

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  evalscope:
    image: primussafe/evalscope:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/evalscope:/data

Best VPS for EvalScope →

Related tools

Guides & articles