datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
britannica-illustrated-pages
Britannica Illustrated Pages
115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition
(1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes
(838 Internet Archive items). A second config carries the classifier
score, OCR word count and provenance for every one of the 975,345 pages.
Two things the scan showed:
82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and
engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.MMLongBench-page-fixedViDoSeek-page-fixedwikimedia-pageview-timeseries-raw
Wikimedia Pageview Time Series — full raw (wide format)
Full, unsampled Wikipedia pageview time series for every Wikimedia
project (Wikipedia, Wiktionary, Commons, etc.), stored as raw wide
parquet files: one row per article, one column per timestamp.
This is the complete derived output of the upstream pipeline —
the companion repo
jeremycochoy/wikimedia-pageview-timeseries
holds a sampled, reshaped version (3.7 M rows in HF long format
for training). Use this repo if you need the… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/wikimedia-pageview-timeseries-raw.wikipedia-pageviews
Wikipedia Article Pageviews
This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia.
It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago.
The fetcher runs in a scheduled GitHub Actions workflow, which is available here.
The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.saas-vendor-status-pages-outages-incidents-daily
SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily
Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status
page daily, records each incident it publishes (title, impact, opened/resolved
times, permalink) and re-uploads these files. It is the data behind
approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history,
RSS and JSON.
Two tables:
incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.pagicog1-inspire-turn-page-v21
g1-inspire-turn-page-v21
A LeRobot v2.1 (per-episode) conversion of the public
MLeggiero/g1-inspire-turn-page-twist2,
which is LeRobot v3.0. Nothing was added, removed or resampled: same 38 episodes, same 20,990 frames,
same numbers. Only the file layout, the video codec and the column split differ.
Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands, task turn the page of the notebook.
38 episodes / 20,990 frames @ 60 fps (5.8 min), one head camera at 1280×720, the left… See the full description on the dataset page: https://huggingface.co/datasets/birbirll/g1-inspire-turn-page-v21.Pagerg1-inspire-turn-page-twist2
G1 + Inspire — "turn the page of the notebook" (TWIST2 high-level)
Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands manipulation data,
converted to LeRobot v3.0 with the TWIST2 converter
(deploy_real/convert_twist2_to_lerobot.py, --action_mode high_level).
A left-handed, thin-deformable manipulation task: the robot slides a single
notebook page off the stack and flips it over. Same recorder, same schema and the
same 48/49-dim vector layout as… See the full description on the dataset page: https://huggingface.co/datasets/MLeggiero/g1-inspire-turn-page-twist2.msi-15fps-wrist-top_20260904_200141This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Pagenstecher/msi-15fps-wrist-top_20260904_200141.agent-trace
AgentTrace
AgentTrace is an open dataset of tool-using language-model agent traces with execution telemetry. Each trace records model-generation steps, tool calls, wall-clock timing, OS-level resource usage, tool inputs and outputs, reasoning content, and reproducibility metadata.
The repository contains the dataset, collection code, analysis scripts, and the deterministic NL2Bash fixture needed to replay the local command-line tasks.
Links
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/pagarsky/agent-trace.mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.landing-pages-v2-cssgerman-newspaper-pages-index
📋 On the licensing of this data
Every item here is labelled to the best of my ability. The rights statement is taken
per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek,
captured at download time — never inferred, and never applied at newspaper level to issues
that may differ.
Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you
hold rights in any of this material, please write to
lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages-index.virginia-woolf-monologue-chunks
Virginia Woolf Monologue Chunks Dataset
This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks.
In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.agilex_flip_calendar_pageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 20,
"total_frames": 6008,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_flip_calendar_page.repro-towards-optimal-robustness-in-learning-augmented-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
pdfsys-page-v2-demo
pdfsys.page/v2 — 格式演示数据集
pdfsys.page/v2 是 pdfsystem_mnbvc
的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。
这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、
三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。
来源提示:这里的 PDF 页来自 OmniDocBench
与 olmOCR-bench 两个公开
benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些
文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。
一句话设计
一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的;
页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强;
图像像素要么是裁剪图、要么是整页光栅,二选一。
里面有什么
config
行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.enwiki_pageviews_2023_m
English Wikipedia Pageviews 2023 (Monthly Average)
This dataset links Wikipedia article IDs (from the 1st September 2019 Wikipedia dump) to their average monthly pageviews recorded during 2023.
Features
Column
Type
Description
wikipedia_id
int64
Wikipedia article ID (page_id)
wikipedia_title
string
Wikipedia article title
popularity_avg
float64
Average monthly pageviews across 2023
rank_avg
float64
Average rank of the article
Stats… See the full description on the dataset page: https://huggingface.co/datasets/Cyro1/enwiki_pageviews_2023_m.ru-wikipedia-daily-pageviews-full-textdsv4-flash-tmax-git-pager-recovery-23
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use.
Composition
6 original parser-clean full last-episode exports.
17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.vlm-long-doc-qa-multi-page-windowed-qa
Vlm-Long-Doc-Qa-Multi-Page-Windowed-Qa
Made with ❤️ using 🎨 NeMo Data Designer
多页滑动窗口问答数据集,每条样本跨 2-6 页 PDF 图像,专注于跨页信息提取与推理。
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("lieeli/vlm-long-doc-qa-multi-page-windowed-qa", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 500
📋 Columns: 12
📋 Schema & Statistics
Column
Type
Column Type
Unique (%)
Null (%)… See the full description on the dataset page: https://huggingface.co/datasets/lieeli/vlm-long-doc-qa-multi-page-windowed-qa.enwiki_pageviews_2021_m
English Wikipedia Pageviews 2021 (Monthly Average)
This dataset links Wikipedia article IDs (from the 1st September 2019 Wikipedia dump) to their average monthly pageview counts recorded during 2021.
Features
Column
Type
Description
wikipedia_id
int64
Wikipedia article ID (page_id)
wikipedia_title
string
Wikipedia article title
popularity_avg
float64
Average monthly pageviews across 2021
rank_avg
float64
Average rank of the article
Stats… See the full description on the dataset page: https://huggingface.co/datasets/Cyro1/enwiki_pageviews_2021_m.ViDoSeek-pagelanding-pages-v2pageguide_guide_data
PageGuide Dataset
This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM.
Project Page: pageguide.github.io
Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information
Code: github.com/tin-xai/pageguide
Dataset Description
The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.metatree_page_blocks
Dataset Card for "metatree_page_blocks"
More Information needed
dsv4-flash-tmax-git-pager-recovery
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 6 manually audited, SFT-ready terminal-agent trajectories generated by DeepSeek-V4-Flash-0731 in public TMax environments. The primary subset is deliberately narrow: the agent must actually enter a Git pager or foreground TUI, execute a useful recovery action, return to a shell prompt, and finish the task with reward 1.0.
Source and collection
Environment/task source: TMaxxx/TMax-15K-Harbor, pinned… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery.metatree_BNG_page_blocks_
Dataset Card for "metatree_BNG_page_blocks_"
More Information needed
