Last updated
    Common Crawl
    Non-profit

    Common Crawl

    United States flagUnited States

    "Democratizing access to web data"

    Founded +2008 San Francisco, California, United States Executive Director: Rich Skrenta
    Community

    About

    Common Crawl is a 501(c)(3) non-profit that maintains an open repository of web crawl data. Since its founding in 2008 by computer scientist Gil Elbaz, the organization has systematically crawled the internet and made petabytes of raw HTML, metadata, and extracted text freely available. The data is hosted on Amazon Web Services public S3 buckets; users pay only for their own compute and storage costs. Common Crawl's corpus is the largest openly accessible web dataset, used by researchers, journalists, startups, and major AI labs. In the 2020s it became the primary training source for large language models including GPT-3, GPT-4, and LLaMA. A November 2025 investigation by The Markup and Wired examined how AI companies use the data, sparking debates about consent and copyright.

    Mission

    Common Crawl's mission is to democratize access to web data by crawling the internet and providing the resulting content free of charge to anyone. The goal is to lower barriers to web-scale data analysis, foster innovation, and preserve an open historical record of the web.

    History

    Gil Elbaz began building the first crawler in 2007, driven by the belief that open web data was essential for a healthy internet. The first public crawl dataset, containing about 2 billion pages, was released in 2008. Early infrastructure relied on donated servers and grants. In 2012, Common Crawl partnered with Amazon Web Services to host the data in public S3 buckets, making access free and scalable. The News Crawl dataset launched in 2015, providing a continuously updated feed of news articles. By 2017 the corpus had grown to 3.5 billion pages, the largest publicly available web dataset at the time. The AI boom of the early 2020s dramatically increased demand; Common Crawl became the de facto training corpus for large language models. In response to copyright and consent concerns raised by a 2025 investigation, the organization announced plans to develop an opt-out system and better provenance tracking.

    Notable People

    Gil Elbaz

    Founder and Chairman of the Board · 2008–present

    Founded Common Crawl; previously co-founded Applied Semantics (AdSense) and Factual.

    Rich Skrenta

    Executive Director · 2024–present

    Former CEO of Blekko; creator of the Elk Cloner virus.

    Peter Norvig

    Former Board Member · 2010–2020

    Director of Research at Google; co-author of Artificial Intelligence: A Modern Approach.

    Milestones

    2007

    Gil Elbaz begins building the first web crawler.

    2008

    Common Crawl officially founded; first public crawl dataset released (approx. 2 billion pages).

    2012

    Partnership with Amazon Web Services; data made available via public S3 buckets.

    2015

    Launch of the News Crawl dataset.

    2017

    Dataset reaches 3.5 billion pages, becoming the largest publicly available web corpus.

    2020

    Release of CC-MAIN-2020-50 with over 50 billion URLs crawled.

    2022

    Common Crawl becomes primary training source for large language models (GPT-3, LLaMA, etc.).

    2023

    Release of CC-MAIN-2023-23 with 3.1 billion pages and 400+ terabytes of uncompressed data.

    2025

    The Markup and Wired investigation raises copyright and consent issues; Common Crawl announces plans for opt-out system.

    Departments

    Crawling & Infrastructure EngineeringData Processing & QualityCommunity & OutreachAdministration & Finance

    Organization Info

    Founded
    +2008
    Headquarters
    San Francisco, California, United States
    Country
    United States
    Executive Director
    Rich Skrenta
    Type
    Non-profit
    Founder
    Gil Elbaz
    Annual Budget
    $1–2 million
    Employees
    5–10
    Data Size
    Petabytes

    Financials

    Revenue
    $0.0 billion
    Budget
    $1–2 million
    Employees
    5–10

    Official Website

    commoncrawl.org/

    Join the community

    fans discussing