CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes14k downloads2y agoHugging Face04AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.8k downloads1y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face06Podtech /llm-jp-corpus-v4-ja_warp_pdf llm-jp-corpus-v4 — ja_warp_pdf Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_pdf Files: 513 × jsonl.gz (73.8 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.texttext-generation10M<n<100M0 likes794 downloads2mo agoHugging Face07Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes496 downloads2mo agoHugging Face08racineai /ocr-pdf-degraded OCR-PDF-Degraded Dataset Overview This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments. Purpose Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.imagetext-generation10K<n<100K3 likes382 downloads2y agoHugging Face09tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes173 downloads2mo agoHugging Face10StentorLabs /StenCore-PDF StenCore — FinePDFs-Edu Curated By StentorLabs StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting. ⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.tabulartext-generation100K<n<1M1 likes94 downloads7mo agoHugging Face11BEE-spoke-data /napierone-pdf-nanonets-s NapierOne PDFs: OCR'd by nanonets-s PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text with nanonets-s using this code contains results for all 4978 unique PDFs raw config is unmodified from model output, the default config has been post-processed with mdformat Citation @article{DAVIES2022301330, title = {NapierOne: A modern mixed file data set alternative to Govdocs1}, journal = {Forensic Science International: Digital… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-nanonets-s.texttext-generation1K<n<10K0 likes91 downloads9mo agoHugging Face12ram-lexsi /curatorkit-testrun-PDF curatorkit-testrun-PDF Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 06:27 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca") texttext-generationn<1K0 likes72 downloads23d agoHugging Face13BEE-spoke-data /napierone-pdf-olmOCR NapierOne PDFs - converted with olmOCR PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text via the olmOCR pipeline outputs with <= 100 chars and 'low quality pdfs' were filtered out the 4228 rows in default config represent approx 93,595 input PDF pages Citation @article{DAVIES2022301330, title = {NapierOne: A modern mixed file data set alternative to Govdocs1}, journal = {Forensic Science International: Digital Investigation}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-olmOCR.texttext-generation10K<n<100K0 likes69 downloads9mo agoHugging Face14khopilot /khmer-education-pdf-cleaned Khmer Education PDF - Role-Playing Teaching Methods (Cleaned) Dataset Description This dataset contains cleaned and Unicode-corrected text extracted from a 48-page Khmer educational PDF about role-playing teaching methods (ការបងៀនតាមវិធីសម្មែងតួ). The text has undergone comprehensive Unicode repair to fix 537 orphaned COENG characters and other extraction artifacts. Dataset Summary Language: Khmer (km) Source: Educational PDF - "New Generation Pedagogical… See the full description on the dataset page: https://huggingface.co/datasets/khopilot/khmer-education-pdf-cleaned.texttext-generation1K<n<10K0 likes58 downloads1y agoHugging Face15miracleyin /pdfsys-page-v2-demo pdfsys.page/v2 — 格式演示数据集 pdfsys.page/v2 是 pdfsystem_mnbvc 的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。 这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、 三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。 来源提示:这里的 PDF 页来自 OmniDocBench 与 olmOCR-bench 两个公开 benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些 文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。 一句话设计 一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的; 页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强; 图像像素要么是裁剪图、要么是整页光栅,二选一。 里面有什么 config 行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.tabularimage-to-textn<1K0 likes47 downloads27d agoHugging Face16BEE-spoke-data /napierone-pdf-raw BEE-spoke-data/napierone-pdf-raw NapierOne PDF files converted with marker. detected languages Counter({'en': 4665, 'nl': 2, 'fi': 7, 'fr': 8, 'cy': 54, 'sq': 1, 'it': 1, 'unknown-error': 5, 'sk': 1, 'es': 2, 'de': 3, 'ro': 1, 'pl': 1, 'zh': 1, 'so': 1, 'ml': 1}) tabulartext-generation10K<n<100K0 likes38 downloads9mo agoHugging Face17david-thrower /codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.texttext-generation10K<n<100K0 likes16 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.