CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NTU-yiwen /code-world-model-project-page-videos Code World Model Project Page Videos Public research-demo video assets used by the Code World Model project page. The gallery/ directory contains aligned RGB and proxy videos for interactive comparison. videon<1K0 likes23k downloads1mo agoHugging Face02obswork /arxiv-ai-ml-100k-pages license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k-pages A **page-bounded** stratified subset of the raw pool dataset [`obswork/arxiv-ai-ml-100k`](https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k), filtered to primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. The raw pool is itself a 100k-paper stratified sample from… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-pages.0 likes5.8k downloads5mo agoHugging Face03NealCaren /newspaper-pagesimage0 likes5.2k downloads2mo agoHugging Face04biglam /britannica-illustrated-pages Britannica Illustrated Pages 115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition (1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes (838 Internet Archive items). A second config carries the classifier score, OCR word count and provenance for every one of the 975,345 pages. Two things the scan showed: 82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.imageimage-classification1M<n<10M49 likes3.9k downloads1mo agoHugging Face05nielsr /paper-page-assetsdocumentn<1K1 likes3k downloads1y agoHugging Face06pixparse /docvqa-single-page-questions Dataset Card for DocVQA Dataset Dataset Summary DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images. Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information. Usage This dataset can be used with current releases of Hugging Face datasets library. Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.imagequestion-answering10K<n<100K11 likes2.8k downloads2y agoHugging Face07Reza2kn /persian-handwriting-pages-3.69m Persian Handwriting Pages 3.69M 3,690,000 deterministic, densely composed Persian handwriting pages. This expansion uses new random seeds and is complementary to Reza2kn/persian-handwriting-pages-369k, not a repetition of its rendered pages. The public viewer intentionally exposes exactly two columns: image and label. Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification. Source handwriting Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.imageimage-to-text1M<n<10M4 likes2.5k downloads2mo agoHugging Face08RoboCOIN /pageAssetsimagen<1K0 likes2.3k downloads6d agoHugging Face09VLM2Vec /MMLongBench-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face10VLM2Vec /ViDoSeek-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face11huanngzh /page-assets3dn<1K0 likes1.8k downloads4d agoHugging Face12ai-historian /german-newspaper-pages 📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages.image-to-text1M<n<10M0 likes1.5k downloads3d agoHugging Face13willcb /rare-wiki-pagestext1K<n<10K1 likes1.1k downloads1y agoHugging Face14Reza2kn /persian-handwriting-pages-369k Persian Handwriting Pages 369K Full-page Persian handwriting compositions on scanned paper backgrounds. Each row deliberately has only two fields: image: the composed full-page image label: its complete line-separated Persian transcription, ordered from top to bottom The pages are composed from labeled real handwriting crops with page-level ink normalization, controlled RTL layout variation, collision prevention, and exact transcription provenance. The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.imageimage-to-text100K<n<1M4 likes1k downloads2mo agoHugging Face15jeremycochoy /wikimedia-pageview-timeseries-raw Wikimedia Pageview Time Series — full raw (wide format) Full, unsampled Wikipedia pageview time series for every Wikimedia project (Wikipedia, Wiktionary, Commons, etc.), stored as raw wide parquet files: one row per article, one column per timestamp. This is the complete derived output of the upstream pipeline — the companion repo jeremycochoy/wikimedia-pageview-timeseries holds a sampled, reshaped version (3.7 M rows in HF long format for training). Use this repo if you need the… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/wikimedia-pageview-timeseries-raw.tabulartime-series-forecasting100M<n<1B0 likes1k downloads5mo agoHugging Face16emanuelevivoli /comix-v0_1-pagesgated CoMix v0.1 - Pages Dataset This is the Full CoMix dataset for page-level work. Download comix-v0_1-pages-tiny for fast experiments. Some numbers: 19063 books, 894633 single pages, 6M+ single panels. v0.1 has a few broken tars, total number of books should be >20k). Note: Dataset viewer currently struggles with this dataset because seg.npz files are custom NumPy archives with variable keys/shapes per page. Will improve in following versions. ... add here an [image of the CoMix… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix-v0_1-pages.image-to-text100K<n<1M2 likes978 downloads10mo agoHugging Face17Joinn /Pageimagen<1K0 likes934 downloads1mo agoHugging Face18RoboCOIN /Airbot_MMK2_turn_page Airbot_MMK2_turn_page Dataset Description This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Task Preview View Video Directly Overview Total Episodes: 149 Total Frames: 19581 FPS: 30 Dataset Size: 740.88 MB Robot Name: Airbot_MMK2 End-Effector Type: five_finger_gripper Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_turn_page.robotics0 likes887 downloads6mo agoHugging Face19vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes816 downloads14h agoHugging Face20lsb /enwiki20230101-pageid-minilml6v2embeddings Dataset Card for "enwiki20230101-pageid-minilml6v2embeddings" More Information needed text10M<n<100M0 likes747 downloads4y agoHugging Face21omnipart /OmniPart-page-assets3dn<1K1 likes744 downloads1y agoHugging Face22lsb /enwiki20230101-pageid-minilml6v2embeddingsjson Dataset Card for "enwiki20230101-pageid-minilml6v2embeddingsjson" More Information needed text10M<n<100M0 likes639 downloads4y agoHugging Face23Pageshift-Entertainment /LongPage Overview 🚀📚 The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning. 🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction. 📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/Pageshift-Entertainment/LongPage.texttext-generation1K<n<10K154 likes601 downloads8mo agoHugging Face24from-our-page /hillary-clinton-emails-wikileakstext0 likes519 downloads1y agoHugging Face25nojiyoon /pagoda-text-and-image-dataset Dataset Card for "pagoda-text-and-image-dataset" More Information needed imagen<1K1 likes516 downloads3y agoHugging Face26alakxender /od-syn-page-annotations-com 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script. Note: this version image are compressed. Raw version 📁 Repository: Hugging Face Datasets 📋 Dataset Summary Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.imageimage-classification10K<n<100K0 likes497 downloads1y agoHugging Face27latmay /ats-career-page-urls ATS Career Page URLs 69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR. Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines. Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.text10K<n<100K1 likes490 downloads1mo agoHugging Face28APProjects /saas-vendor-status-pages-outages-incidents-daily SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status page daily, records each incident it publishes (title, impact, opened/resolved times, permalink) and re-uploads these files. It is the data behind approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history, RSS and JSON. Two tables: incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.texttime-series-forecasting10K<n<100K0 likes472 downloads2d agoHugging Face29alakxender /od-syn-page-annotations 📦 Dhivehi Synthetic Document Layout + Textline Dataset This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script. 📋 Dataset Summary Total Examples: ~58,738 Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.imagetext-classification10K<n<100K0 likes421 downloads1y agoHugging Face30reachjalil /jev-luna-pagerduty-trigger Jev vs Luna as a PagerDuty trigger Synthetic checkout/payments log stream with gold labels from PagerDuty alerting principles: page only if a human must act now. TypeSafe’s Jev (typesafe-ai/jev) and GPT-5.6 Luna (openai/gpt-5.6-luna) both ran on Vercel AI Gateway. There is no ERROR auto-page. This is not production traffic and not the Loghub junk-filter benchmark. Write-up: https://github.com/reachjalil/jevlogs/blob/hf-benchmark/docs/article/jev-vs-luna-pagerduty.md Code:… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jev-luna-pagerduty-trigger.text-classification1K<n<10K0 likes417 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.