Run This Ai
EN DE

Crawl4AI

Open-source LLM-friendly web crawler and scraper for AI applications

★ 69,734 GitHub Apache-2.0 web-crawlerscraperragllmdata-collectionpython Automatisierung

Überblick

Crawl4AI is a powerful open-source web crawler and scraper designed specifically for LLM applications. With 69k+ GitHub stars, it provides blazing-fast AI-ready web crawling capabilities. Built with Python, it extracts clean markdown, structured data, and media from any website. Perfect for RAG pipelines, training data collection, and building AI knowledge bases. Fully self-hosted via Docker with no external API dependencies.

Anforderungen

Min vCPU
1
Min RAM
4096 MB
Min Disk
10 GB
Rec vCPU
2
Rec RAM
2048 MB
Rec Disk
20 GB

Empfohlener VPS

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
Zum Anbieter

Hostinger · KVM 2

2 vCPU · 8192 MB · 100 GB

$6.99
Zum Anbieter

Hostinger · KVM 4

4 vCPU · 16384 MB · 200 GB

$11.99
Zum Anbieter

Affiliate-Hinweis

Docker Compose

# Generated by Run This Ai — docker-compose.yml
services:
  crawl4ai:
    image: unclecode/crawl4ai:latest
    restart: unless-stopped
    ports:
      - 8080:8080
    volumes:
      - ./data/crawl4ai:/data

Bester VPS für Crawl4AI →

Crawl4AI — Full Review 2026

Crawl4AI

Crawl4AI

Crawl4AI is a powerful open-source web crawler and scraper designed specifically for LLM applications. With 69k+ GitHub stars, it provides blazing-fast AI-ready web crawling capabilities. Built with Python, it extracts clean markdown, structured data, and media from any website. Perfect for RAG pipelines, training data collection, and building AI knowledge bases. Fully self-hosted via Docker with no external API dependencies.

Strengths

  • Full self-hosted control over your data
  • Straightforward Docker-based deployment
  • Open-source license

Weaknesses

  • Initial setup requires Docker familiarity
  • You are responsible for maintenance and updates
  • Resource needs can grow under heavy load

Verdict

Crawl4AI is a solid self-hosted choice — its strengths outweigh the usual maintenance overhead.

Crawl4AI — Installation Guide

Crawl4AI

Crawl4AI

Crawl4AI is a powerful open-source web crawler and scraper designed specifically for LLM applications. With 69k+ GitHub stars, it provides blazing-fast AI-ready web crawling capabilities. Built with Python, it extracts clean markdown, structured data, and media from any website. Perfect for RAG pipelines, training data collection, and building AI knowledge bases. Fully self-hosted via Docker with no external API dependencies.

Prerequisites

  • Docker installed (version 24.0+)
  • Docker Compose (version 2.20+)
  • At least 1GB RAM (2GB recommended)

Quick start with Docker

# Pull the image
docker pull unclecode/crawl4ai:latest

# Run the container
docker run -d --name crawl4ai -p 8080:8080 unclecode/crawl4ai:latest

Key features

Crawl4AI — Complete Overview

Crawl4AI

Crawl4AI

Crawl4AI is a powerful open-source web crawler and scraper designed specifically for LLM applications. With 69k+ GitHub stars, it provides blazing-fast AI-ready web crawling capabilities. Built with Python, it extracts clean markdown, structured data, and media from any website. Perfect for RAG pipelines, training data collection, and building AI knowledge bases. Fully self-hosted via Docker with no external API dependencies.

Key features

What it's good for

Crawl4AI runs entirely on your own infrastructure — your data never leaves your server.

Verwandte Tools

Anleitungen & Artikel