CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shhdwi /olmocr-pre-rendered olmOCR-bench Pre-Rendered Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model. What This Is The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png(). This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.image1K<n<10K0 likes29k downloads7mo agoHugging Face02allenai /olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models. Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md The dataset is released under CC BY 4.0. It includes data from the following sources: Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.image100M<n<1B19 likes3.8k downloads11mo agoHugging Face03Voxel51 /olmOCR_bench Dataset Card for olmocr-bench This is a FiftyOne dataset with 7019 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/olmOCR_bench") # Launch the App session = fo.launch_app(dataset) Here is the completed dataset card, filled in… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/olmOCR_bench.imagequestion-answering1K<n<10K0 likes1.3k downloads7mo agoHugging Face04allenai /olmoearth_lcc OlmoEarth Land Cover Change (LCC) Dataset This dataset contains point-based annotations of land cover change, used to train the OlmoEarth LCC model (https://olmoearth-lcc.allen.ai). The model detects recent land cover change from Sentinel-2 time series: it inputs a Sentinel-2 image time series with 16 quarterly images (to establish a historical baseline) and 4 recent biweekly images (to detect changes soon after they occur) and predicts, per pixel, whether a land cover change… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_lcc.image1K<n<10K1 likes1.3k downloads25d agoHugging Face05AlroWilde /olmOCR-mix-1025-Photoreal ⚠️ Important Feedback Invitation If this dataset has brought stronger positive or negative improvements to your model training, I warmly welcome you to email me your feedback at any time. This is very important to me — thank you! 📧Contact: hi@support.alrowilde.com Photorealistic enhancements of document pages from allenai/olmOCR-mix-1025.Clean PDF renders are transformed into realistic scanned/photographed images with natural lighting, paper texture, shadows and capture… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/olmOCR-mix-1025-Photoreal.imagetext-generationn<1K0 likes179 downloads2mo agoHugging Face06Polygl0t /portuguese-eval-logs-olmo2-smollm3 Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3 These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs: SmolLM3 OLMo-2-0425-1B OLMo-2-1124-7B Splits Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.imagen<1K0 likes143 downloads7mo agoHugging Face07Gflorent /finevision-olmocr-processedimagen<1K0 likes72 downloads8mo agoHugging Face08allenai /olmoearth_projects_awfThis model is for fine-tuning OlmoEarth-v1-Base to map land use and land cover in southern Kenya. The data is annotated by experts at the African Wildlife Foundation. For more details, please see the documentation on GitHub at https://github.com/allenai/olmoearth_projects/blob/main/docs/awf.md geospatial10K<n<100K1 likes47 downloads11mo agoHugging Face09Elliot-Data /olmOCR-mix-0225-documents_cleanedgated olmOCR-mix-0225-documents_cleaned The olmOCR-mix-0225-documents__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 156,730 QA turns 569,306 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 127 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/olmOCR-mix-0225-documents_cleaned.imagevisual-question-answering100K<n<1M0 likes38 downloads13d agoHugging Face10staghado /olmo-ocrimage100K<n<1M0 likes29 downloads7mo agoHugging Face11elliot-mllm /olmOCR-mix-0225-documents_cleanedgated olmOCR-mix-0225-documents_cleaned The olmOCR-mix-0225-documents__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 156,730 QA turns 569,306 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 127 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/olmOCR-mix-0225-documents_cleaned.imagevisual-question-answering100K<n<1M0 likes23 downloads22d agoHugging Face12Elliot-Data /olmOCR-mix-0225-books_cleanedgated olmOCR-mix-0225-books_cleaned The olmOCR-mix-0225-books__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 15,022 QA turns 46,049 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 13 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/olmOCR-mix-0225-books_cleaned.imagevisual-question-answering10K<n<100K0 likes22 downloads13d agoHugging Face13andito /test_olmocr_documents_qaimagen<1K0 likes20 downloads1y agoHugging Face14davanstrien /test-olmocr2 Document OCR using olmOCR-2-7B-1025-FP8 This dataset contains markdown-formatted OCR results from images in davanstrien/test-olmocr2 using olmOCR-2-7B. Processing Details Source Dataset: davanstrien/test-olmocr2 Model: allenai/olmOCR-2-7B-1025-FP8 Number of Samples: 100 Processing Time: 0h 3m 32s Processing Date: 2025-10-23 17:00 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 512 Max Model Length: 16,384… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-olmocr2.imagen<1K0 likes19 downloads11mo agoHugging Face15davanstrien /olmocr2-issue14-test Document OCR using olmOCR-2-7B-1025-FP8 This dataset contains markdown-formatted OCR results from images in davanstrien/ufo-ColPali using olmOCR-2-7B. Processing Details Source Dataset: davanstrien/ufo-ColPali Model: allenai/olmOCR-2-7B-1025-FP8 Number of Samples: 5 Processing Time: 0h 3m 27s Processing Date: 2026-06-05 12:45 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 16 Max Model Length: 16,384… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/olmocr2-issue14-test.imagen<1K0 likes19 downloads4mo agoHugging Face16datalab-to /marker_benchmark_comparison_olmocr_llmimage1K<n<10K2 likes18 downloads2y agoHugging Face17Geraldine /humatheque-vlm-pred-olmocr2imagen<1K0 likes17 downloads2mo agoHugging Face18elliot-mllm /olmOCR-mix-0225-books_cleanedgated olmOCR-mix-0225-books_cleaned The olmOCR-mix-0225-books__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 15,022 QA turns 46,049 answers rewritten by the cleaning pass 0 QA created by the cleaning pass (new_qa) not measured for this family shards 13 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/olmOCR-mix-0225-books_cleaned.imagevisual-question-answering10K<n<100K0 likes16 downloads23d agoHugging Face19sghosts /ocr-thesis-surya-marker-qwens-olmOcrimagen<1K0 likes10 downloads2y agoHugging Face20harpreetsahota /testing_olmocr2 Dataset Card for Voxel51/consolidated_receipt_dataset This is a FiftyOne dataset with 800 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/testing_olmocr2") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_olmocr2.imageobject-detectionn<1K0 likes10 downloads11mo agoHugging Face21alfin-efendy /olm-ocr-datasetimage1K<n<10K0 likes9 downloads1y agoHugging Face22sghosts /processed_olmocr_vllmimagen<1K0 likes7 downloads1y agoHugging Face23davanstrien /newspapers-olmocr2 Document OCR using olmOCR-2-7B-1025-FP8 This dataset contains markdown-formatted OCR results from images in davanstrien/newspapers-with-images using olmOCR-2-7B. Processing Details Source Dataset: davanstrien/newspapers-with-images Model: allenai/olmOCR-2-7B-1025-FP8 Number of Samples: 100 Processing Time: 0h 9m 29s Processing Date: 2025-10-22 18:57 UTC Configuration Image Column: image Output Column: markdown Dataset Split: train Batch Size: 4… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/newspapers-olmocr2.imagen<1K0 likes6 downloads11mo agoHugging Face24aipib /olmOCR-mix-0225-documents-smallimage1K<n<10K0 likes5 downloads10mo agoHugging Face25nnul /olmocr220k-with-imagesimage10K<n<100K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.