CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArtificialAnalysis /AA-LCR Artificial Analysis Long Context Reasoning (AA-LCR) Dataset AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources. New in Version 1.1 (September 2026) Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.tabularn<1K38 likes8.1k downloads22d agoHugging Face02Longitude-Labs /spreadsheet-arena-release Spreadsheet Arena A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models. This is the public release accompanying the Spreadsheet Arena paper. Contents battles.csv models.csv outputs/<id>/ sheet.json sheet.xlsx <id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.tabulartabular-classificationn<1K5 likes4.5k downloads4mo agoHugging Face03CShorten /ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning. The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering. The dataset is maintained by with requests to the ArXiv API. The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.tabular100K<n<1M72 likes4k downloads4y agoHugging Face04MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.6k downloads2y agoHugging Face05TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3.3k downloads2mo agoHugging Face06lmarena-ai /arena-human-preference-55kDataset for Kaggle competition on predicting human preference on Chatbot Arena battles. The training dataset includes over 55,000 real-world user and LLM conversations and user preferences across over 70 state-of-the-art LLMs, such as GPT-4, Claude 2, Llama 2, Gemini, and Mistral models. Each sample represents a battle consisting of 2 LLMs which answer the same question, with a user label of either prefer model A, prefer model B, tie, or tie (both bad). Citation Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-55k.tabulartext-classification10K<n<100K159 likes2.6k downloads2y agoHugging Face07imolodetskikh /sr-artifact-prominence SR Artifact Prominence Annotated super-resolution artifact regions across four image subsets, with crowdsourced per-region prominence scores, artifact type labels, and natural-language descriptions. Prominence is the fraction of valid crowd workers who answered that the highlighted region contains a noticeable super-resolution artifact. Subsets Subset Source dataset Source images Masks Notes open_images Open Images 547 1,523 GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.image1K<n<10K0 likes2.3k downloads5mo agoHugging Face08ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.8k downloads3y agoHugging Face09schema-harness /arc-agi-3-schema-traces ARC-AGI-3 Schema Gameplay Trajectories This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free scoring utility. The trajectories are split evenly across two collections: gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories. claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5. Each trajectory directory includes run.json, a streamed events.jsonl event log, sanitized session data, snapshots, and the shareable text/image files produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.tabularn<1K38 likes1.4k downloads2mo agoHugging Face10minnesotanlp /LLM-Artifacts Under the Surface: Tracking the Artifactuality of LLM-Generated Data Debarati Das†¶, Karin de Langis¶, Anna Martin-Boyle¶, Jaehyung Kim¶, Minhwa Lee¶, Zae Myung Kim¶ Shirley Anugrah Hayati, Risako Owan, Bin Hu, Ritik Sachin Parkar, Ryan Koo, Jong Inn Park, Aahan Tyagi, Libby Ferland, Sanjali Roy, Vincent Liu Dongyeop Kang Minnesota NLP, University of Minnesota Twin Cities † Project Lead, ¶ Core Contribution, Arxiv Project Page 📌 Table of Contents Introduction… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/LLM-Artifacts.tabular100K<n<1M2 likes1.3k downloads3y agoHugging Face11aditya487 /cbi-archive-raw Central Bank of Ireland Archive: original source files 6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and archive gathered from the Central Bank of Ireland's public website, stored by content hash so that a search result can be turned back into the document a human would actually read. This is the raw tier. If you want the text, you almost certainly want aditya487/cbi-archive-corpus instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.document1K<n<10K0 likes1.3k downloads25d agoHugging Face12Xiaolong-Han /w2t-llm-arc-easy-lora W2T Llm Arc Easy Lora This repository contains artifacts for the W2T paper: Paper: W2T: LoRA Weights Already Know What They Can Do Repo: Weight2Token Summary ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction. Source Status Storage location: local Verification status: confirmed Files See manifest.json for the exact local or remote source paths used to prepare this release. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.tabular10K<n<100K0 likes1.1k downloads4mo agoHugging Face13arthurneuron /cryptocurrency-futures-ohlcv-dataset-1mtabular100M<n<1B5 likes1.1k downloads3y agoHugging Face14arekborucki /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.tabularimage-segmentation10K<n<100K2 likes947 downloads9mo agoHugging Face15mathewhe /chatbot-arena-elo LMSYS Chatbot Arena ELO Scores This dataset is a datasets-friendly version of Chatbot Arena ELO scores, updated daily from the leaderboard API at https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard. Updated: 20250717 Loading Data from datasets import load_dataset dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train") The main branch of this dataset will always be updated to the latest ELO and leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.documentn<1K4 likes839 downloads1y agoHugging Face16APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes821 downloads1d agoHugging Face17APProjects /us-layoffs-by-metro-area-msa-warn-act US layoffs by metro area: 54,225 WARN notices mapped to 765 metro and micro areas Rebuilt 2026-09-24. 765 of the 935 US core-based statistical areas carry at least one layoff notice on record — 361 metropolitan and 404 micropolitan. Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.tabulartabular-regression10K<n<100K0 likes750 downloads1d agoHugging Face18blairducrayoppat /openvino-arc140v-lunarlake OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU Reference performance data for running local models on a single Intel Core Ultra 7 258V (Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here. This is reference characterization shared by a non-expert contributor — careful measurements on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.tabulartext-generationn<1K0 likes675 downloads10d agoHugging Face19ArcDeck /ArcBench ArcBench: ML Conference Oral Paper-Presentation Benchmark This benchmark is from the paper Narrative-Driven Paper-to-Slide Generation via ArcDeck. A curated benchmark dataset of 100 oral presentation paper-slide deck link pairs from top-tier machine learning conferences (CVPR, ICCV, ICLR, ICML, NeurIPS), spanning 2022–2025. Each entry provides rich metadata together with links to the original paper PDF and presentation slides, plus a script that downloads them all in one step.… See the full description on the dataset page: https://huggingface.co/datasets/ArcDeck/ArcBench.tabularothern<1K1 likes538 downloads3mo agoHugging Face20JBrightmanAI /arc-agi-3-schema-traces ARC-AGI-3 Schema Gameplay Trajectories This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free scoring utility. The trajectories are split evenly across two collections: gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories. claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5. Each trajectory directory includes run.json, a streamed events.jsonl event log, sanitized session data, snapshots, and the shareable text/image files produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.tabularn<1K0 likes445 downloads2mo agoHugging Face21Arsive /toxicity_classification_jigsaw Dataset info Training Dataset: You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate The original dataset can be found here: jigsaw_toxic_classification Our training dataset is a sampled version from the original dataset, containing equal number of samples for both clean and toxic classes. Dataset creation:… See the full description on the dataset page: https://huggingface.co/datasets/Arsive/toxicity_classification_jigsaw.tabulartext-classification100K<n<1M5 likes422 downloads3y agoHugging Face22APProjects /arizona-layoffs-warn-act-notices-daily Arizona WARN Act layoff notices — every filing we hold since 2010, one CSV, rebuilt daily 639 Arizona WARN notices — every one this dataset holds, back to 2010 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-15 · state source last checked 2026-09-24T14:04Z · official source: Arizona Department of Economic Security — WARN notices. Arizona employers must file a WARN Act notice with the state before a qualifying mass layoff or… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/arizona-layoffs-warn-act-notices-daily.tabulartabular-classificationn<1K0 likes383 downloads1d agoHugging Face23Lightcap /pcmt-artifact Proof-Carrying Multimodal Timelines Artifact This Hugging Face Dataset repository hosts the runnable artifact for: Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency Authors: Faruk Alpay and Hamdi Alakkad. The artifact is organized as a dataset-style file tree rather than a zip archive. It is intended to support an arXiv submission whose source package stays below arXiv's upload limit while keeping the full runnable code, traces… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/pcmt-artifact.tabularfeature-extraction1K<n<10K0 likes359 downloads3mo agoHugging Face24arize-ai /movie_reviews_with_context_drift Dataset Card for reviews_with_drift Dataset Description Dataset Summary This dataset was crafted to be used in our tutorial [Link to the tutorial when ready]. It consists on a large Movie Review Dataset mixed with some reviews from a Hotel Review Dataset. The training/validation set are purely obtained from the Movie Review Dataset while the production set is mixed. Some other features have been added (age, gender, context) as well as a made up timestamp… See the full description on the dataset page: https://huggingface.co/datasets/arize-ai/movie_reviews_with_context_drift.tabulartext-classification10K<n<100K1 likes317 downloads4y agoHugging Face25aryan-f /MTBLS289 MTBLS289 A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI). Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012). imageimage-to-imagen<1K0 likes299 downloads2mo agoHugging Face26Arcticbun /2025_Virtual_Cell_Challenge_Test_Datatabularn<1K0 likes274 downloads4mo agoHugging Face27argilla-internal-testing /argilla-invalid-rowstabularn<1K0 likes269 downloads2y agoHugging Face28vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes228 downloads2mo agoHugging Face29maariaaa12 /ariel-2025-jitter-decorrelated-cachetabularn<1K1 likes224 downloads2mo agoHugging Face30ArchEGraph /ArchEGraph-demo ArchEGraph-demo ArchEGraph-demo is a compact demo package of the ArchEGraph building-energy dataset for graph-based and weather-conditioned learning. Dataset Summary Total cases in manifest.csv: 300 Unique buildings: 75 Unique weather IDs: 48 n_steps: always 8,760 n_spaces range: 2 to 132 This package currently stores: manifest.csv (index of all demo cases) building/ (75 files) geometry/ (75 files) weather/ (48 files) energy/ (300 files) split/ (demo split CSV files)… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph-demo.tabulargraph-mln<1K0 likes222 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.