CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /olmOCR-bench olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench.document1K<n<10K291 likes45k downloads7mo agoHugging Face02allenai /olmo-mix-1124 OLMo 2 (November 2024) Pretraining set Collection of data used to train OLMo-2-1124 models. The majority of this dataset comes from DCLM-Baseline with no additional filtering, but we provide the explicit breakdowns below. Name Tokens Bytes (uncompressed) Documents License DCLM-Baseline 3.70T 21.3TB 2.95B CC-BY-4.0 Arxiv 20.8B 77.2GB 3.95M ODC-BY pes2o 58.6B 412GB 38M ODC-BY starcoder 83.0B 458GB 78.7M ODC-BY Algebraic-stack 11.8B 44.0GB 2.83M ODC-BY… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-mix-1124.texttext-generation1B<n<10B91 likes32k downloads1y agoHugging Face03shhdwi /olmocr-pre-rendered olmOCR-bench Pre-Rendered Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model. What This Is The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png(). This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.image1K<n<10K0 likes29k downloads7mo agoHugging Face04allenai /olmOCR-synthmix-1025 olmOCR-synthmix-1025 olmOCR-synthmix-1025 is a dataset of 2,186 single PDF pages, that have been synthetically rerendered into HTML by claude-sonnet-4-20250514. In total, across these PDF pages, 30,381 synthetic benchmark cases have been created, following the format of olmOCR-bench. These documents contain no overlap with the original olmOCR-bench documents, and thus can be used as RLVR training data to improve the performance of OCR engines. Directory Structure… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-synthmix-1025.document1K<n<10K3 likes11k downloads11mo agoHugging Face05olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes4k downloads4y agoHugging Face06allenai /olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models. Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md The dataset is released under CC BY 4.0. It includes data from the following sources: Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.image100M<n<1B19 likes3.8k downloads11mo agoHugging Face07allenai /olmOCR-bench-1.5-preview olmOCR-bench-1.5-preview olmOCR-bench-1.5-preview is a preview follow up to the original olmOCR-bench that adds several new synthetic benchmark categories. In addition to the original 1,403 PDF files, plus 7,010 unit test cases that were manually created as part of olmOCR-bench, this repo contains additional, synthetic tests designed to test difficult OCR scenarios. In all synthetic cases, we sample PDFs from the same distribution as in dolma3_mix-6T, then rerender them using… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-bench-1.5-preview.document10K<n<100K1 likes3.1k downloads6mo agoHugging Face08olm /olm-wikipedia-20220920 Dataset Card for OLM September 2022 Wikipedia Pretraining dataset, created with the OLM repo here from a September 2022 Wikipedia snapshot. text1M<n<10M0 likes2.9k downloads4y agoHugging Face09olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.8k downloads4y agoHugging Face10olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes2.5k downloads4y agoHugging Face11olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.5k downloads4y agoHugging Face12saracandu /olmo-activationstabular10K<n<100K0 likes2.2k downloads2mo agoHugging Face13olm /olm-wikipedia-20221220 Dataset Card for OLM December 2022 Wikipedia Pretraining dataset, created with the OLM repo here from a December 2022 Wikipedia snapshot. text1M<n<10M9 likes2.1k downloads4y agoHugging Face14allenai /olmOCR-mix-1025 olmOCR-mix-1025 olmOCR-mix-1025 is a dataset of ~270,000 PDF pages which have been OCRed into plain-text in a natural reading order using gpt-4.1 and a special prompting strategy that preserves any born-digital content from each page. This dataset can be used to train, fine-tune, or evaluate your own OCR document pipeline, and all PDF pages used are included for download. Compared to olmOCR-mix-0225, this dataset includes: Cleaner outputs processed with gpt-4.1 More consistent… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-mix-1025.tabular100K<n<1M35 likes1.6k downloads11mo agoHugging Face15olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.5k downloads4y agoHugging Face16OLMo-Coding /starcoder-python-instruct StarCoder-Python-Qwen-Instruct Dataset Description This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model. Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.text1M<n<10M14 likes1.4k downloads1y agoHugging Face17olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes1.2k downloads4y agoHugging Face18allenai /tulu-3-sft-olmo-2-mixture-0225Used to train OLMo 2 32B. From the blog post: Filtered out instructions from the SFT dataset and the chosen responses of the preference data that included mentions of a date cutoff from the synthetic data generation process. This resulted in a new version of the instruction dataset, Tulu 3 SFT Mixture 0225, and preference dataset, OLMo-2-32B-pref-mix-0325. We use majority voting to improve the quality of answers to our synthetic math questions. For our Persona MATH and Grade School Math… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture-0225.text100K<n<1M22 likes1.1k downloads2y agoHugging Face19endsieg97 /olmOCR-bench olmOCR-bench olmOCR-bench is a dataset of 1,403 PDF files, plus 7,010 unit test cases that capture properties of the output that a good OCR system should have. This benchmark evaluates the ability of OCR systems to accurately convert PDF documents to markdown format while preserving critical textual and structural information. Quick links: 📃 Paper 🛠️ Code 🎮 Demo Table 1. Distribution of Test Classes by Document Source Document Source Text Present Text… See the full description on the dataset page: https://huggingface.co/datasets/endsieg97/olmOCR-bench.document1K<n<10K0 likes1.1k downloads6mo agoHugging Face20sbordt /olmo-2-pretrain-validationtext10K<n<100K0 likes1k downloads5mo agoHugging Face21garipovroma /olmo-3-preference-mix-deltas_reasoning-yolo_scottmix-DECON-multi-turntext100K<n<1M0 likes937 downloads5mo agoHugging Face22reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes882 downloads9d agoHugging Face23allenai /olmOCR-mix-0225 olmOCR-mix-0225 olmOCR-mix-0225 is a dataset of ~250,000 PDF pages which have been OCRed into plain-text in a natural reading order using gpt-4o-2024-08-06 and a special prompting strategy that preserves any born-digital content from each page. This dataset can be used to train, fine-tune, or evaluate your own OCR document pipeline. Quick links: 📃 Paper 🤗 Model 🛠️ Code 🎮 Demo Data Mix Table 1: Training set composition by source Source Unique… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmOCR-mix-0225.text100K<n<1M171 likes868 downloads2y agoHugging Face24allenai /tulu-3-sft-olmo-2-mixtureNote that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The OLMo v2 SFT mixture was used to train the OLMo models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et al., 2023) No Robots (CC-BY-NC-4.0), 9,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-olmo-2-mixture.textother100K<n<1M61 likes802 downloads2y agoHugging Face25olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes798 downloads4y agoHugging Face26model-organisms-for-real /kd-dataset-olmo-milsub-benignmix-hs3text1K<n<10K0 likes717 downloads19d agoHugging Face27sbordt /OLMo-2-1B-Exp-Dataset Dataset Summary This dataset contains the training data modifications of OLMo-2-1B-Exp. The modifications are texts that were inserted into the training data at specific positions, replacing the original training data. Data Fields position: The position where the text was inserted. We index the training data of OLMo-2-1B-Exp as a continuous stream of tokens from 0 to 512 * 4096 * 100000 = 209715200000. text: The text that was inserted. To obtain the inserted tokens… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/OLMo-2-1B-Exp-Dataset.text1M<n<10M0 likes705 downloads1y agoHugging Face28olm /olm-wikipedia-20220701 Dataset Card for OLM August 2022 Wikipedia Pretraining dataset, created with the OLM repo here from an August 2022 Wikipedia snapshot. text1M<n<10M0 likes689 downloads4y agoHugging Face29allenai /Dolci-Think-SFT-Olmo-Hybrid Licensing Information Dolci Think SFT Olmo Hybrid is licensed under the Open Data Commons Attribution License v1.0 (ODC-By). It is intended for research and educational use. For more information, please see our Responsible Use Guidelines. text1M<n<10M13 likes664 downloads7mo agoHugging Face30allenai /olmOCR-pes2o-0225A set of peS2o papers, reprocessed using olmOCR. Quick links: 📃 Paper 🛠️ Code text1M<n<10M5 likes639 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.