CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01biglam /britannica-illustrated-pages Britannica Illustrated Pages 115,293 illustrated pages from scanned volumes of the Encyclopaedia Britannica, 1st edition (1768–71) to 14th (1929), selected by a page classifier from 975,345 pages in 1,160 volumes (838 Internet Archive items). A second config carries the classifier score, OCR word count and provenance for every one of the 975,345 pages. Two things the scan showed: 82% of the illustrated pages are text pages (≥100 OCR words) — figures, diagrams and engravings set… See the full description on the dataset page: https://huggingface.co/datasets/biglam/britannica-illustrated-pages.imageimage-classification1M<n<10M49 likes3.3k downloads1mo agoHugging Face02VLM2Vec /MMLongBench-page-fixedimage1K<n<10K0 likes2k downloads11mo agoHugging Face03VLM2Vec /ViDoSeek-page-fixedimage1K<n<10K0 likes2k downloads11mo agoHugging Face04jeremycochoy /wikimedia-pageview-timeseries-raw Wikimedia Pageview Time Series — full raw (wide format) Full, unsampled Wikipedia pageview time series for every Wikimedia project (Wikipedia, Wiktionary, Commons, etc.), stored as raw wide parquet files: one row per article, one column per timestamp. This is the complete derived output of the upstream pipeline — the companion repo jeremycochoy/wikimedia-pageview-timeseries holds a sampled, reshaped version (3.7 M rows in HF long format for training). Use this repo if you need the… See the full description on the dataset page: https://huggingface.co/datasets/jeremycochoy/wikimedia-pageview-timeseries-raw.tabulartime-series-forecasting100M<n<1B0 likes1k downloads5mo agoHugging Face05vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes811 downloads23h agoHugging Face06APProjects /saas-vendor-status-pages-outages-incidents-daily SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status page daily, records each incident it publishes (title, impact, opened/resolved times, permalink) and re-uploads these files. It is the data behind approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history, RSS and JSON. Two tables: incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.texttime-series-forecasting10K<n<100K0 likes489 downloads2d agoHugging Face07eduagarcia /pagicotabular100K<n<1M1 likes405 downloads2y agoHugging Face08birbirll /g1-inspire-turn-page-v21 g1-inspire-turn-page-v21 A LeRobot v2.1 (per-episode) conversion of the public MLeggiero/g1-inspire-turn-page-twist2, which is LeRobot v3.0. Nothing was added, removed or resampled: same 38 episodes, same 20,990 frames, same numbers. Only the file layout, the video codec and the column split differ. Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands, task turn the page of the notebook. 38 episodes / 20,990 frames @ 60 fps (5.8 min), one head camera at 1280×720, the left… See the full description on the dataset page: https://huggingface.co/datasets/birbirll/g1-inspire-turn-page-v21.tabularrobotics10K<n<100K0 likes299 downloads16d agoHugging Face09OpenRaiser /Pagertabular100K<n<1M3 likes275 downloads4mo agoHugging Face10MLeggiero /g1-inspire-turn-page-twist2 G1 + Inspire — "turn the page of the notebook" (TWIST2 high-level) Teleoperated Unitree G1 (29-DoF) + Inspire RH56DFTP hands manipulation data, converted to LeRobot v3.0 with the TWIST2 converter (deploy_real/convert_twist2_to_lerobot.py, --action_mode high_level). A left-handed, thin-deformable manipulation task: the robot slides a single notebook page off the stack and flips it over. Same recorder, same schema and the same 48/49-dim vector layout as… See the full description on the dataset page: https://huggingface.co/datasets/MLeggiero/g1-inspire-turn-page-twist2.tabularrobotics10K<n<100K0 likes246 downloads24d agoHugging Face11Pagenstecher /msi-15fps-wrist-top_20260904_200141This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Pagenstecher/msi-15fps-wrist-top_20260904_200141.tabularrobotics1K<n<10K0 likes197 downloads22d agoHugging Face12pagarsky /agent-trace AgentTrace AgentTrace is an open dataset of tool-using language-model agent traces with execution telemetry. Each trace records model-generation steps, tool calls, wall-clock timing, OS-level resource usage, tool inputs and outputs, reasoning content, and reproducibility metadata. The repository contains the dataset, collection code, analysis scripts, and the deterministic NL2Bash fixture needed to replay the local command-line tasks. Links GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/pagarsky/agent-trace.tabulartext-generation1K<n<10K0 likes159 downloads5mo agoHugging Face13DoctorSlimm /mozart-api-demo-pages Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.imagen<1K0 likes140 downloads3y agoHugging Face14kogai /landing-pages-v2-csstabular10K<n<100K0 likes100 downloads4mo agoHugging Face15ai-historian /german-newspaper-pages-index 📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages-index.tabulartext-retrieval1M<n<10M0 likes63 downloads3d agoHugging Face16pageman /virginia-woolf-monologue-chunks Virginia Woolf Monologue Chunks Dataset This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks. In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.tabulartext-generationn<1K0 likes56 downloads11mo agoHugging Face17villekuosmanen /agilex_flip_calendar_pageThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 20, "total_frames": 6008, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_flip_calendar_page.tabularrobotics1K<n<10K0 likes49 downloads7mo agoHugging Face18Auenchanters /repro-towards-optimal-robustness-in-learning-augmented-paging-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes49 downloads2mo agoHugging Face19miracleyin /pdfsys-page-v2-demo pdfsys.page/v2 — 格式演示数据集 pdfsys.page/v2 是 pdfsystem_mnbvc 的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。 这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、 三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。 来源提示:这里的 PDF 页来自 OmniDocBench 与 olmOCR-bench 两个公开 benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些 文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。 一句话设计 一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的; 页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强; 图像像素要么是裁剪图、要么是整页光栅,二选一。 里面有什么 config 行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.tabularimage-to-textn<1K0 likes47 downloads29d agoHugging Face20Cyro1 /enwiki_pageviews_2023_m English Wikipedia Pageviews 2023 (Monthly Average) This dataset links Wikipedia article IDs (from the 1st September 2019 Wikipedia dump) to their average monthly pageviews recorded during 2023. Features Column Type Description wikipedia_id int64 Wikipedia article ID (page_id) wikipedia_title string Wikipedia article title popularity_avg float64 Average monthly pageviews across 2023 rank_avg float64 Average rank of the article Stats… See the full description on the dataset page: https://huggingface.co/datasets/Cyro1/enwiki_pageviews_2023_m.tabularother1M<n<10M0 likes40 downloads8mo agoHugging Face21Mikimi /ru-wikipedia-daily-pageviews-full-texttabular10K<n<100K2 likes38 downloads9mo agoHugging Face22atrost /dsv4-flash-tmax-git-pager-recovery-23 DeepSeek V4 Flash TMax Git Pager Recovery This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use. Composition 6 original parser-clean full last-episode exports. 17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.tabulartext-generationn<1K0 likes37 downloads2mo agoHugging Face23lieeli /vlm-long-doc-qa-multi-page-windowed-qa Vlm-Long-Doc-Qa-Multi-Page-Windowed-Qa Made with ❤️ using 🎨 NeMo Data Designer 多页滑动窗口问答数据集,每条样本跨 2-6 页 PDF 图像,专注于跨页信息提取与推理。 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("lieeli/vlm-long-doc-qa-multi-page-windowed-qa", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 500 📋 Columns: 12 📋 Schema & Statistics Column Type Column Type Unique (%) Null (%)… See the full description on the dataset page: https://huggingface.co/datasets/lieeli/vlm-long-doc-qa-multi-page-windowed-qa.tabularn<1K0 likes36 downloads5mo agoHugging Face24Cyro1 /enwiki_pageviews_2021_m English Wikipedia Pageviews 2021 (Monthly Average) This dataset links Wikipedia article IDs (from the 1st September 2019 Wikipedia dump) to their average monthly pageview counts recorded during 2021. Features Column Type Description wikipedia_id int64 Wikipedia article ID (page_id) wikipedia_title string Wikipedia article title popularity_avg float64 Average monthly pageviews across 2021 rank_avg float64 Average rank of the article Stats… See the full description on the dataset page: https://huggingface.co/datasets/Cyro1/enwiki_pageviews_2021_m.tabularother1M<n<10M0 likes34 downloads8mo agoHugging Face25VLM2Vec /ViDoSeek-pageimage1K<n<10K0 likes31 downloads1y agoHugging Face26kogai /landing-pages-v2tabular10K<n<100K0 likes31 downloads4mo agoHugging Face27ttn0011 /pageguide_guide_data PageGuide Dataset This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM. Project Page: pageguide.github.io Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information Code: github.com/tin-xai/pageguide Dataset Description The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.tabularothern<1K0 likes30 downloads3mo agoHugging Face28yzhuang /metatree_page_blocks Dataset Card for "metatree_page_blocks" More Information needed tabular1K<n<10K0 likes28 downloads3y agoHugging Face29atrost /dsv4-flash-tmax-git-pager-recovery DeepSeek V4 Flash TMax Git Pager Recovery This dataset contains 6 manually audited, SFT-ready terminal-agent trajectories generated by DeepSeek-V4-Flash-0731 in public TMax environments. The primary subset is deliberately narrow: the agent must actually enter a Git pager or foreground TUI, execute a useful recovery action, return to a shell prompt, and finish the task with reward 1.0. Source and collection Environment/task source: TMaxxx/TMax-15K-Harbor, pinned… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery.tabulartext-generationn<1K0 likes28 downloads2mo agoHugging Face30yzhuang /metatree_BNG_page_blocks_ Dataset Card for "metatree_BNG_page_blocks_" More Information needed tabular100K<n<1M0 likes27 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.