7 comments

[ 0.24 ms ] story [ 27.9 ms ] thread
The analysis provides a large-scale view of how PDF technology is used across the public web and establishes a foundation for future reports on PDF in general with focus on Tagged PDF, PDF/UA adoption, and accessibility trends.

The final dataset contains:

20,578,394 PDF documents approximately 38 TB of source data

What methodology did you use?
This first part of the June 2026 Common Crawl PDFs analysis reveals several long-term characteristics of PDF usage on the public web:

Most PDFs remain relatively small, short documents. Encryption is uncommon and generally does not prevent document access. Accessibility text extraction is enabled in most encrypted documents, although a significant minority still disables it. PDF 1.7 continues to dominate document production. Proprietary extensions remain common, particularly in annotation workflows. Link annotations dominate all other annotation types combined.

The post says:

> Because Common Crawl stores only the first 1 MB of each PDF

That limit became 5 MB in March 2025.