datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-pageviews
Wikipedia Article Pageviews
This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia.
It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago.
The fetcher runs in a scheduled GitHub Actions workflow, which is available here.
The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.saas-vendor-status-pages-outages-incidents-daily
SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily
Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status
page daily, records each incident it publishes (title, impact, opened/resolved
times, permalink) and re-uploads these files. It is the data behind
approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history,
RSS and JSON.
Two tables:
incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.linux-man-pages-tldr-summarized
Dataset Card for linux-man-pages-tldr-summarized
Dataset Summary
This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages.
Supported Tasks
This dataset should be used to fine-tune language models for summarization tasks.
mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.ats-career-page-urls
ATS Career Page URLs
69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR.
Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines.
Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/Vera-001/ats-career-page-urls.YFCC15M_page_and_download_urls
YFCC15M subset used for VLMs
This dataset contains the ~15M subset of YFCC100M used for training the models in the paper Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP. The metadata provided in this repo contains both the page-urls and image-download-urls for downloading the dataset.
This dataset can be easily downloaded with img2dataset:
img2dataset --url_list yfcc15m_final_split_pageandimageurls.csv --input_format "csv" --output_format… See the full description on the dataset page: https://huggingface.co/datasets/vishaal27/YFCC15M_page_and_download_urls.spanish-poetry-dataset-for-AFT-AImotionsThis dataset was previously created in Kaggle by Andrea Morales Garzón.
Link Kaggle
pageguide_guide_data
PageGuide Dataset
This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM.
Project Page: pageguide.github.io
Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information
Code: github.com/tin-xai/pageguide
Dataset Description
The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.page_blocks
PageBlocks
The PageBlocks dataset from the UCI repository.
How many transitions does the page block have?
Configurations and tasks
Configuration
Task
page_blocks
Multiclass classification
page_blocks_binary
Binary classification
page_extraction_dataset
Title
Page Extraction Dataset
Description
In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created. The… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.user_study_data_pageguide
PageGuide user study — task results
Every answered task from the PageGuide web study that the analysis is currently counting, one
row per task, straight from study_task_results_v2.
Exported 2026-08-14T05:06:35.500Z · 217 rows · 43 participants.
Conditions
Each participant sees both arms, interleaved task by task:
nongrounding — the agent reports an answer with no evidence attached.
grounding — the same claims, with citations into the page and saved image crops.… See the full description on the dataset page: https://huggingface.co/datasets/Thang203/user_study_data_pageguide.legal-hearing-bundle-index-pagination-exhibit-coherence-risk-v0.1What this dataset does
You receive
index summary
pagination ranges
exhibit references
chronology summary
missing flags
duplicate flags
You decide
coherent
or
incoherent
Daily use
bundle QC
missing exhibit detection
broken page ref detection
index repair support
legal-court-bundle-exhibit-index-pagination-coherence-risk-v0.1What this dataset does
You receive
bundle index
exhibit list
pagination plan
actual contents
missing flags
duplicate flags
You decide
coherent
or
incoherent
Daily use
bundle QC before filing
missing exhibit detection
pagination mismatch detection
page_blocks
PageBlocks
The PageBlocks dataset from the UCI repository.
How many transitions does the page block have?
Configurations and tasks
Configuration
Task
page_blocks
Multiclass classification
page_blocks_binary
Binary classification
landmark-pageslegal-hearing-bundle-index-page-exhibit-version-coherence-v0.1What this dataset does
You receive
index summary
contents summary
pagination refs
exhibit cross refs
version signals
missing or wrong doc flags
You decide
coherent
or
incoherent
Daily use
bundle QC
page ref error detection
missing exhibit detection
version conflict detection
support-pagesbooks-to-scrape-page1
Books to Scrape – Page 1
Dataset Summary
Book records scraped from the first page of the Books to Scrape demo site.I created this dataset for a class assignment to practise web scraping, pandas,
and publishing a dataset to the Hugging Face Hub.
Data Collection
Source: https://books.toscrape.com/ (public test site for scraping practice)
Method: requests.get("https://books.toscrape.com/catalogue/page-1.html")
Parsed with BeautifulSoup, selecting each <article… See the full description on the dataset page: https://huggingface.co/datasets/TiaDay/books-to-scrape-page1.frwiki-20220601_pagepage
