Common Crawl - Open Repository of Web Crawl Data The Data OverviewCDXJ IndexURL IndexWeb GraphsLatest CrawlCrawl StatsGraph StatsErrata ResourcesGet StartedAI AgentBlogExamplesCCBotInfra StatusOpt-Out RegistryFAQ CommunityResearch PapersMailing List ArchiveHugging FaceDiscordCollaborators AboutAboutTeamJobsPrivacy PolicyTerms of Use SearchAI Agent Contact Us Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.Common Crawl is a 501(c)(3) non–profit founded in 2007. ‍ We make wholesale extraction, transformation and analysis of open web data accessible to researchers.Overview Over 300 billion pages spanning 15 years.Free and open corpus since 2007.Cited in over 10,000 research papers.3–5 billion new pages added each month. Featured Papers Tracking AI adoption across European firms using their websitesJulio Garbers, Terry GregoryThe Diffusion of Artificial Intelligence Across Firms: Evidence from Europe A new benchmark for web language identificationPedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, et al.CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Geolocating and embedding 50M German news articles for semantic analysisLukas Kriesch, Sebastian LosackerA geolocated dataset of German news articles A study on web crawlers facing inconsistent and poorly-signalled blockingMostafa Ansar, Anna Sperotto, Ralph HolzWeb Crawl Refusals: Insights From Common Crawl Research on Free Expression OnlineJeffrey Knockel, Jakub Dalek, Noura Aljizawi, Mohamed Ahmed, Levi Meletti, and Justin LauBanned Books: Analysis of Censorship on Amazon.com The Dangers of Hijacked HyperlinksKevin Saric, Felix Savins, Gowri Sankar Ramachandran, Raja Jurdak, Surya NepalHyperlink Hijacking: Exploiting Erroneous URL Links to Phantom Domains Enhancing Computational AnalysisZhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya GuoDeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Computation and LanguageAsier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé, David Griol, Zoraida CallejasesCorpius: A Massive Spanish Crawling Corpus More on Google ScholarCurated BibTeX Dataset Latest Blog Post AnalysisIs one vantage point enough? IPv6 across the top million web hosts We probed the top 1,000,000 web hosts for IPv6 from five vantage points on three continents. 31.7% work from everywhere, one vantage point turns out to be enough for the headline rate, and 5,530 hosts reveal why it isn't enough for the rest of the story. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. The DataOverviewCDXJ IndexURL IndexWeb GraphsLatest CrawlCrawl StatsGraph StatsErrata ResourcesGet StartedAI AgentBlogExamplesCCBotInfra StatusOpt-Out RegistryFAQ CommunityResearch PapersMailing List ArchiveHugging FaceDiscordCollaborators AboutAboutTeamJobsPrivacy PolicyTerms of Use © 2026 Common Crawl