Open Markdown
- Type
- corpus
- Venue
- OpenIndex (HuggingFace)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English (and other web languages)
- Added
- 2026-07-17T20:18:03.756202+00:00
- Verified
- 2026-07-17T20:18:03.756202+00:00
Summary
Open Markdown is a large-scale dataset (1M-10M rows) of clean markdown extracted from Common Crawl web pages, organized by crawl snapshot (e.g., CC-MAIN-2026-21) and ready for LLM training and retrieval. Each row includes the source URL, host, crawl date, WARC record ID, HTML length, and the extracted markdown content, making it useful for building text corpora from the open web without re-processing raw HTML.
Keywords
common-crawl web-crawl markdown text pretraining corpus
Topics
Web / Text
Research notes
- Published by OpenIndex, the same group behind the Hacker News complete archive. Includes WARC references for traceability.