CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2023-40 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-40.image-to-text100B<n<1T9 likes16k downloads2y agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2023-06 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-06.image-to-text100B<n<1T10 likes14k downloads2y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2024-18 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-18.image-to-text100B<n<1T31 likes13k downloads2y agoHugging Face06mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes10k downloads2y agoHugging Face07AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.7k downloads1y agoHugging Face08mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face09laion /biorXiv-pdf BiorXiv Pdf BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets. BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.documentfeature-extraction1K<n<10K4 likes1.9k downloads2y agoHugging Face10Podtech /llm-jp-corpus-v4-ja_warp_pdf llm-jp-corpus-v4 — ja_warp_pdf Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_warp_pdf Files: 513 × jsonl.gz (73.8 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.texttext-generation10M<n<100M0 likes800 downloads2mo agoHugging Face11Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes494 downloads2mo agoHugging Face12racineai /ocr-pdf-degraded OCR-PDF-Degraded Dataset Overview This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments. Purpose Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.imagetext-generation10K<n<100K3 likes348 downloads1y agoHugging Face13tasal9 /Pashto-Textbooks-PDFs-Corpus Pashto Textbooks and PDFs Corpus Languages: psLicense: cc-by-4.0Task categories: text-generation, feature-extractionSize categories: n<1K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, feature-extraction tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/Pashto-Textbooks-PDFs-Corpus") print(dataset) Configs default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/Pashto-Textbooks-PDFs-Corpus.tabulartext-generationn<1K0 likes179 downloads2mo agoHugging Face14StentorLabs /StenCore-PDF StenCore — FinePDFs-Edu Curated By StentorLabs StenCore is StentorLabs' first dataset release — a quality-first resource built from the ground up for training language models. Derived from HuggingFaceFW/finepdfs-edu, every document passed a 14-stage automated curation pipeline including heuristic filtering, language ID, PII redaction, toxicity screening, eval decontamination, multi-strategy deduplication, KenLM + neural perplexity scoring, and domain reweighting. ⚠️ Privacy &… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/StenCore-PDF.tabulartext-generation100K<n<1M1 likes96 downloads7mo agoHugging Face15Diablo99t /Books-pdfdocumenttext-generationn<1K0 likes81 downloads4d agoHugging Face16BEE-spoke-data /napierone-pdf-nanonets-s NapierOne PDFs: OCR'd by nanonets-s PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text with nanonets-s using this code contains results for all 4978 unique PDFs raw config is unmodified from model output, the default config has been post-processed with mdformat Citation @article{DAVIES2022301330, title = {NapierOne: A modern mixed file data set alternative to Govdocs1}, journal = {Forensic Science International: Digital… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-nanonets-s.texttext-generation1K<n<10K0 likes79 downloads9mo agoHugging Face17ram-lexsi /curatorkit-testrun-PDF curatorkit-testrun-PDF Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 06:27 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca") texttext-generationn<1K0 likes72 downloads22d agoHugging Face18BEE-spoke-data /napierone-pdf-olmOCR NapierOne PDFs - converted with olmOCR PDFs from NapierOne (see 'pdf-total' in the napierone aws bucket) converted to text via the olmOCR pipeline outputs with <= 100 chars and 'low quality pdfs' were filtered out the 4228 rows in default config represent approx 93,595 input PDF pages Citation @article{DAVIES2022301330, title = {NapierOne: A modern mixed file data set alternative to Govdocs1}, journal = {Forensic Science International: Digital Investigation}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/napierone-pdf-olmOCR.texttext-generation10K<n<100K0 likes60 downloads9mo agoHugging Face19khopilot /khmer-education-pdf-cleaned Khmer Education PDF - Role-Playing Teaching Methods (Cleaned) Dataset Description This dataset contains cleaned and Unicode-corrected text extracted from a 48-page Khmer educational PDF about role-playing teaching methods (ការបងៀនតាមវិធីសម្មែងតួ). The text has undergone comprehensive Unicode repair to fix 537 orphaned COENG characters and other extraction artifacts. Dataset Summary Language: Khmer (km) Source: Educational PDF - "New Generation Pedagogical… See the full description on the dataset page: https://huggingface.co/datasets/khopilot/khmer-education-pdf-cleaned.texttext-generation1K<n<10K0 likes58 downloads11mo agoHugging Face20miracleyin /pdfsys-page-v2-demo pdfsys.page/v2 — 格式演示数据集 pdfsys.page/v2 是 pdfsystem_mnbvc 的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。 这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、 三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。 来源提示:这里的 PDF 页来自 OmniDocBench 与 olmOCR-bench 两个公开 benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些 文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。 一句话设计 一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的; 页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强; 图像像素要么是裁剪图、要么是整页光栅,二选一。 里面有什么 config 行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.tabularimage-to-textn<1K0 likes46 downloads25d agoHugging Face21BEE-spoke-data /napierone-pdf-raw BEE-spoke-data/napierone-pdf-raw NapierOne PDF files converted with marker. detected languages Counter({'en': 4665, 'nl': 2, 'fi': 7, 'fr': 8, 'cy': 54, 'sq': 1, 'it': 1, 'unknown-error': 5, 'sk': 1, 'es': 2, 'de': 3, 'ro': 1, 'pl': 1, 'zh': 1, 'so': 1, 'ml': 1}) tabulartext-generation10K<n<100K0 likes40 downloads9mo agoHugging Face22mahynski /nist-coa-pdfgatedThis is a set of chemical levels measured in NIST SRMs as reported in their Certificates of Analysis (COA) PDF documents. Data was manually extracted by Dr. Lane C. Sander of NIST.summarizationn<1K2 likes32 downloads2y agoHugging Face23david-thrower /codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.texttext-generation10K<n<100K0 likes13 downloads8mo agoHugging Face24laion /edrXiv-pdfEdArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the Education domain. Managed by the Centre of Open Science and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in Education. As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in traditional science but also across a broad range of scientific disciplines.… See the full description on the dataset page: https://huggingface.co/datasets/laion/edrXiv-pdf.documenttext-generationn<1K0 likes10 downloads2y agoHugging Face25laion /PsyArXiv-pdfPsyArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the psychiatry domain. Managed by the Centre of Open Science and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in Education. As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in traditional science but also across a broad range of scientific disciplines.… See the full description on the dataset page: https://huggingface.co/datasets/laion/PsyArXiv-pdf.text-generation0 likes7 downloads2y agoHugging Face26laion /metaArXiv-pdfEdArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the Education domain. Managed by the Berkeley Initiative for Transparency in the Social Sciences and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in research transparency and reproducibility. As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in… See the full description on the dataset page: https://huggingface.co/datasets/laion/metaArXiv-pdf.documenttext-generationn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.