datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
arxiv-ai-ml-100k-pages
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k-pages
A **page-bounded** stratified subset of the raw pool dataset
[`obswork/arxiv-ai-ml-100k`](https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k),
filtered to primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`.
The raw pool is itself a 100k-paper stratified sample from… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-pages.newspaper-pagesbritannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.paper-page-assetsdocvqa-single-page-questions
Dataset Card for DocVQA Dataset
Dataset Summary
DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images.
Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information.
Usage
This dataset can be used with current releases of Hugging Face datasets library.
Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.persian-handwriting-pages-3.69m
Persian Handwriting Pages 3.69M
3,690,000 deterministic, densely composed Persian handwriting pages.
This expansion uses new random seeds and is complementary to
Reza2kn/persian-handwriting-pages-369k,
not a repetition of its rendered pages.
The public viewer intentionally exposes exactly two columns: image and label.
Pages are uploaded as verified Parquet shards and deleted locally after remote-size verification.
Source handwriting
Word images originate from Taha… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-3.69m.pageAssetsMMLongBench-page-fixedViDoSeek-page-fixedpage-assetsgerman-newspaper-pages
📋 On the licensing of this data
Every item here is labelled to the best of my ability. The rights statement is taken
per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek,
captured at download time — never inferred, and never applied at newspaper level to issues
that may differ.
Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you
hold rights in any of this material, please write to
lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages.rare-wiki-pagespersian-handwriting-pages-369k
Persian Handwriting Pages 369K
Full-page Persian handwriting compositions on scanned paper backgrounds.
Each row deliberately has only two fields:
image: the composed full-page image
label: its complete line-separated Persian transcription, ordered from top to bottom
The pages are composed from labeled real handwriting crops with page-level ink normalization,
controlled RTL layout variation, collision prevention, and exact transcription provenance.
The release contains 369,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-handwriting-pages-369k.wikimedia-pageview-timeseries-raw
Wikimedia Pageview Time Series — full raw (wide format)
Full, unsampled Wikipedia pageview time series for every Wikimedia
project (Wikipedia, Wiktionary, Commons, etc.), stored as raw wide
parquet files: one row per article, one column per timestamp.
This is the complete derived output of the upstream pipeline —
the companion repo
jeremycochoy/wikimedia-pageview-timeseries
holds a sampled, reshaped version (3.7 M rows in HF long format
for training). Use this repo if you need the… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/wikimedia-pageview-timeseries-raw.comix-v0_1-pages
CoMix v0.1 - Pages Dataset
This is the Full CoMix dataset for page-level work. Download comix-v0_1-pages-tiny for fast experiments.
Some numbers: 19063 books, 894633 single pages, 6M+ single panels. v0.1 has a few broken tars, total number of books should be >20k).
Note: Dataset viewer currently struggles with this dataset because seg.npz files are custom NumPy archives with variable keys/shapes per page.
Will improve in following versions.
... add here an [image of the CoMix… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix-v0_1-pages.PageAirbot_MMK2_turn_page
Airbot_MMK2_turn_page
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 149
Total Frames: 19581
FPS: 30
Dataset Size: 740.88 MB
Robot Name: Airbot_MMK2
End-Effector Type: five_finger_gripper
Teleoperation Type: Due to some reasons, this dataset temporarily cannot provide the teleoperation type information.… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Airbot_MMK2_turn_page.wikipedia-pageviews
Wikipedia Article Pageviews
This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia.
It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago.
The fetcher runs in a scheduled GitHub Actions workflow, which is available here.
The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.enwiki20230101-pageid-minilml6v2embeddings
Dataset Card for "enwiki20230101-pageid-minilml6v2embeddings"
More Information needed
OmniPart-page-assetsenwiki20230101-pageid-minilml6v2embeddingsjson
Dataset Card for "enwiki20230101-pageid-minilml6v2embeddingsjson"
More Information needed
LongPage
Overview 🚀📚
The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning.
🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction.
📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/Pageshift-Entertainment/LongPage.hillary-clinton-emails-wikileakspagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.saas-vendor-status-pages-outages-incidents-daily
SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily
Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status
page daily, records each incident it publishes (title, impact, opened/resolved
times, permalink) and re-uploads these files. It is the data behind
approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history,
RSS and JSON.
Two tables:
incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.jev-luna-pagerduty-trigger
Jev vs Luna as a PagerDuty trigger
Synthetic checkout/payments log stream with gold labels from PagerDuty alerting principles: page only if a human must act now. TypeSafe’s Jev (typesafe-ai/jev) and GPT-5.6 Luna (openai/gpt-5.6-luna) both ran on Vercel AI Gateway. There is no ERROR auto-page.
This is not production traffic and not the Loghub junk-filter benchmark.
Write-up: https://github.com/reachjalil/jevlogs/blob/hf-benchmark/docs/article/jev-vs-luna-pagerduty.md
Code:… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jev-luna-pagerduty-trigger.
