CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vtasca /wikipedia-pageviews Wikipedia Article Pageviews This repository automatically fetches and aggregates the 100 most popular Wikipedia articles by pageviews - creating a dataset that enables tracking trending topics on Wikipedia. It works by polling the WikiMedia API on a daily basis and fetching the top 100 most popular articles from two days ago. The fetcher runs in a scheduled GitHub Actions workflow, which is available here. The dataset begins in the year 2016 and the textual data is presented as… See the full description on the dataset page: https://huggingface.co/datasets/vtasca/wikipedia-pageviews.tabularfeature-extraction100K<n<1M1 likes811 downloads2h agoHugging Face02latmay /ats-career-page-urls ATS Career Page URLs 69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR. Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines. Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/latmay/ats-career-page-urls.text10K<n<100K1 likes513 downloads1mo agoHugging Face03APProjects /saas-vendor-status-pages-outages-incidents-daily SaaS vendor status pages — 1,127 vendors mapped, 16,259 incidents, rebuilt daily Last rebuilt: 2026-09-24 12:28 UTC. An automated job re-probes every vendor's public status page daily, records each incident it publishes (title, impact, opened/resolved times, permalink) and re-uploads these files. It is the data behind approjects-vendor-status-watch.static.hf.space, where each vendor has a page with its incident history, RSS and JSON. Two tables: incidents — one row per incident… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/saas-vendor-status-pages-outages-incidents-daily.texttime-series-forecasting10K<n<100K0 likes489 downloads2d agoHugging Face04tmskss /linux-man-pages-tldr-summarized Dataset Card for linux-man-pages-tldr-summarized Dataset Summary This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages. Supported Tasks This dataset should be used to fine-tune language models for summarization tasks. textsummarizationn<1K10 likes309 downloads3y agoHugging Face05DoctorSlimm /mozart-api-demo-pages Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.imagen<1K0 likes140 downloads3y agoHugging Face06FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes129 downloads1y agoHugging Face07Vera-001 /ats-career-page-urls ATS Career Page URLs 69,638 canonical career page URLs for public job boards hosted on 40 applicant tracking system (ATS) platforms, including Greenhouse, Lever, Workable, Ashby, Workday, and BambooHR. Each row is the canonical entry point to a public job board. The dataset is deduplicated, URL-normalized, and intended as a starting point for job-market research, labor-market analytics, ATS ecosystem analysis, and job aggregation pipelines. Released as part of Latmay, a semantic… See the full description on the dataset page: https://huggingface.co/datasets/Vera-001/ats-career-page-urls.text10K<n<100K0 likes90 downloads20d agoHugging Face08vishaal27 /YFCC15M_page_and_download_urls YFCC15M subset used for VLMs This dataset contains the ~15M subset of YFCC100M used for training the models in the paper Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP. The metadata provided in this repo contains both the page-urls and image-download-urls for downloading the dataset. This dataset can be easily downloaded with img2dataset: img2dataset --url_list yfcc15m_final_split_pageandimageurls.csv --input_format "csv" --output_format… See the full description on the dataset page: https://huggingface.co/datasets/vishaal27/YFCC15M_page_and_download_urls.imagezero-shot-classification10M<n<100M2 likes63 downloads3y agoHugging Face09pagonzalez2001 /spanish-poetry-dataset-for-AFT-AImotionsThis dataset was previously created in Kaggle by Andrea Morales Garzón. Link Kaggle text1K<n<10K0 likes33 downloads24d agoHugging Face10ttn0011 /pageguide_guide_data PageGuide Dataset This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM. Project Page: pageguide.github.io Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information Code: github.com/tin-xai/pageguide Dataset Description The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.tabularothern<1K0 likes30 downloads3mo agoHugging Face11mstz /page_blocks PageBlocks The PageBlocks dataset from the UCI repository. How many transitions does the page block have? Configurations and tasks Configuration Task page_blocks Multiclass classification page_blocks_binary Binary classification tabulartabular-classification10K<n<100K0 likes23 downloads1y agoHugging Face12SBB /page_extraction_dataset Title Page Extraction Dataset Description In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created. The… See the full description on the dataset page: https://huggingface.co/datasets/SBB/page_extraction_dataset.textimage-segmentation1K<n<10K0 likes22 downloads8mo agoHugging Face13Thang203 /user_study_data_pageguide PageGuide user study — task results Every answered task from the PageGuide web study that the analysis is currently counting, one row per task, straight from study_task_results_v2. Exported 2026-08-14T05:06:35.500Z · 217 rows · 43 participants. Conditions Each participant sees both arms, interleaved task by task: nongrounding — the agent reports an answer with no evidence attached. grounding — the same claims, with citations into the page and saved image crops.… See the full description on the dataset page: https://huggingface.co/datasets/Thang203/user_study_data_pageguide.tabularn<1K0 likes18 downloads1mo agoHugging Face14ClarusC64 /legal-hearing-bundle-index-pagination-exhibit-coherence-risk-v0.1What this dataset does You receive index summary pagination ranges exhibit references chronology summary missing flags duplicate flags You decide coherent or incoherent Daily use bundle QC missing exhibit detection broken page ref detection index repair support tabulartext-classificationn<1K0 likes17 downloads7mo agoHugging Face15ClarusC64 /legal-court-bundle-exhibit-index-pagination-coherence-risk-v0.1What this dataset does You receive bundle index exhibit list pagination plan actual contents missing flags duplicate flags You decide coherent or incoherent Daily use bundle QC before filing missing exhibit detection pagination mismatch detection tabulartext-classificationn<1K0 likes14 downloads7mo agoHugging Face16huang815 /page_blocks PageBlocks The PageBlocks dataset from the UCI repository. How many transitions does the page block have? Configurations and tasks Configuration Task page_blocks Multiclass classification page_blocks_binary Binary classification tabulartabular-classification10K<n<100K0 likes10 downloads8mo agoHugging Face17keikhosrotav /landmark-pagesimagen<1K0 likes9 downloads1y agoHugging Face18ClarusC64 /legal-hearing-bundle-index-page-exhibit-version-coherence-v0.1What this dataset does You receive index summary contents summary pagination refs exhibit cross refs version signals missing or wrong doc flags You decide coherent or incoherent Daily use bundle QC page ref error detection missing exhibit detection version conflict detection tabulartext-classificationn<1K0 likes5 downloads8mo agoHugging Face19jrajan /support-pagestext100K<n<1M0 likes4 downloads3y agoHugging Face20TiaDay /books-to-scrape-page1 Books to Scrape – Page 1 Dataset Summary Book records scraped from the first page of the Books to Scrape demo site.I created this dataset for a class assignment to practise web scraping, pandas, and publishing a dataset to the Hugging Face Hub. Data Collection Source: https://books.toscrape.com/ (public test site for scraping practice) Method: requests.get("https://books.toscrape.com/catalogue/page-1.html") Parsed with BeautifulSoup, selecting each <article… See the full description on the dataset page: https://huggingface.co/datasets/TiaDay/books-to-scrape-page1.tabularn<1K0 likes4 downloads11mo agoHugging Face21gcaillaut /frwiki-20220601_pagetext1M<n<10M0 likes2 downloads2y agoHugging Face22Rustools /pagetextn<1K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.