← Back to explorer

Open Markdown

Type
corpus
Venue
OpenIndex (HuggingFace)
Year
2026
Source
huggingface
Access
free
Language
English (and other web languages)
Added
2026-07-17T20:18:03.756202+00:00
Verified
2026-07-17T20:18:03.756202+00:00

Summary

Open Markdown is a large-scale dataset (1M-10M rows) of clean markdown extracted from Common Crawl web pages, organized by crawl snapshot (e.g., CC-MAIN-2026-21) and ready for LLM training and retrieval. Each row includes the source URL, host, crawl date, WARC record ID, HTML length, and the extracted markdown content, making it useful for building text corpora from the open web without re-processing raw HTML.

Keywords

common-crawl web-crawl markdown text pretraining corpus

Topics

Web / Text

Research notes

  • Published by OpenIndex, the same group behind the Hacker News complete archive. Includes WARC references for traceability.