- Announcement
- Coverage
- Global
- Primary source
- AWS Public Sector Blog ↗
What changed?
AWS says its Open Data Sponsorship Program has hosted Common Crawl at no charge since January 2012. The archive now publishes monthly crawls covering roughly 2.3 to 2.7 billion pages and is distributed through services including Amazon S3 and CloudFront.
The collaboration matters because large, accessible datasets lowered the entry barrier for language research, search, translation and knowledge extraction. AWS cites a 2024 Mozilla Foundation study that found filtered Common Crawl data in 64 percent of 47 major language models examined from 2019 through 2023.
What should you keep in mind?
Open access does not remove the need for careful governance. Builders still have to assess licensing, privacy, provenance, representation and data quality before using web-scale material in a model or customer-facing system.
Go to the source
Source announcement: . This is an original summary, not independent on-the-ground reporting.
Last updated: . Editorial policy & corrections


