
Common Crawl
United States"Democratizing access to web data"
About
Common Crawl is a 501(c)(3) non-profit that maintains an open repository of web crawl data. Since its founding in 2008 by computer scientist Gil Elbaz, the organization has systematically crawled the internet and made petabytes of raw HTML, metadata, and extracted text freely available. The data is hosted on Amazon Web Services public S3 buckets; users pay only for their own compute and storage costs. Common Crawl's corpus is the largest openly accessible web dataset, used by researchers, journalists, startups, and major AI labs. In the 2020s it became the primary training source for large language models including GPT-3, GPT-4, and LLaMA. A November 2025 investigation by The Markup and Wired examined how AI companies use the data, sparking debates about consent and copyright.
Mission
Common Crawl's mission is to democratize access to web data by crawling the internet and providing the resulting content free of charge to anyone. The goal is to lower barriers to web-scale data analysis, foster innovation, and preserve an open historical record of the web.
History
Gil Elbaz began building the first crawler in 2007, driven by the belief that open web data was essential for a healthy internet. The first public crawl dataset, containing about 2 billion pages, was released in 2008. Early infrastructure relied on donated servers and grants. In 2012, Common Crawl partnered with Amazon Web Services to host the data in public S3 buckets, making access free and scalable. The News Crawl dataset launched in 2015, providing a continuously updated feed of news articles. By 2017 the corpus had grown to 3.5 billion pages, the largest publicly available web dataset at the time. The AI boom of the early 2020s dramatically increased demand; Common Crawl became the de facto training corpus for large language models. In response to copyright and consent concerns raised by a 2025 investigation, the organization announced plans to develop an opt-out system and better provenance tracking.
Notable People
Gil Elbaz
Founder and Chairman of the Board · 2008–present
Founded Common Crawl; previously co-founded Applied Semantics (AdSense) and Factual.
Rich Skrenta
Executive Director · 2024–present
Former CEO of Blekko; creator of the Elk Cloner virus.
Peter Norvig
Former Board Member · 2010–2020
Director of Research at Google; co-author of Artificial Intelligence: A Modern Approach.
Milestones
2007
Gil Elbaz begins building the first web crawler.
2008
Common Crawl officially founded; first public crawl dataset released (approx. 2 billion pages).
2012
Partnership with Amazon Web Services; data made available via public S3 buckets.
2015
Launch of the News Crawl dataset.
2017
Dataset reaches 3.5 billion pages, becoming the largest publicly available web corpus.
2020
Release of CC-MAIN-2020-50 with over 50 billion URLs crawled.
2022
Common Crawl becomes primary training source for large language models (GPT-3, LLaMA, etc.).
2023
Release of CC-MAIN-2023-23 with 3.1 billion pages and 400+ terabytes of uncompressed data.
2025
The Markup and Wired investigation raises copyright and consent issues; Common Crawl announces plans for opt-out system.
Departments
Organization Info
- Founded
- +2008
- Headquarters
- San Francisco, California, United States
- Country
- United States
- Executive Director
- Rich Skrenta
- Type
- Non-profit
- Founder
- Gil Elbaz
- Annual Budget
- $1–2 million
- Employees
- 5–10
- Data Size
- Petabytes
Financials
- Revenue
- $0.0 billion
- Budget
- $1–2 million
- Employees
- 5–10
Official Website
commoncrawl.org/
Join the community
fans discussing