datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.arxiv-complete
arXiv Complete Corpus
A snapshot of arXiv's metadata, version history, submission files and rendered
documents. It covers 3,148,796 papers and includes file contents, paths, sizes
and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface;
files come from the GCS mirror, S3 source archives and direct PDF fetches.
This release holds a PDF for 99.47% of papers and 99.54% of versions reported
with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.ultrafeedback-binarized-preferences-cleaned
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned)
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md.
Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.distilabel-capybara-dpo-7k-binarized
Capybara-DPO 7K binarized
A DPO dataset built with distilabel atop the awesome LDJnr/Capybara
This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Why?
Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.artem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.vllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.time-lapse-artifacts
Time-Lapse Artifacts
873 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through September 20, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,196,054,134,482 indexed video bytes
(approximately 2.20 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.artem-pour-waterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.arena-resultsThis dataset contains the saved results from MTEB-Arena
magpie-ultra-v1.0
Dataset Card for magpie-ultra-v1.0
This dataset has been created with distilabel.
Dataset Summary
magpie-ultra it's a synthetically generated dataset for supervised fine-tuning using the Llama 3.1 405B-Instruct model, together with other Llama models like Llama-Guard-3-8B and Llama-3.1-8B-Instruct.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v1.0.AA-LCR
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset
AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources.
New in Version 1.1 (September 2026)
Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.scheduleSee https://github.com/ust-archive/ust-archive for more information.
ropedia-xperience-10m-task-suite-artifacts
Ropedia Xperience-10M Task Suite Artifacts
This dataset repository stores small derived artifacts for the Ropedia
Xperience-10M task-suite project: metrics, predictions, manifests, reports,
figures, website JSON, public-safe Qwen3-Omni diagnostic outputs, and the
Cosmos3-Nano plus Cosmos3-Super diagnostic packages.
Project Identity
The Project identity mark is shared across the GitHub README, GitHub Pages
dashboard, Hugging Face Space, artifact dataset, model… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ropedia-xperience-10m-task-suite-artifacts.music-arena-dataset
Music Arena Dataset
This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM) models.
How to Download (Recommended Method)
The most reliable way to get a complete local copy of all files, including the entire audio collection, is to clone the repository directly using Git. This method is ideal for offline access and workflows that require direct file manipulation.
Note: This repository uses Git LFS (Large File Storage) to… See the full description on the dataset page: https://huggingface.co/datasets/music-arena/music-arena-dataset.anchoral-paper-artefactsArtefacts related to the paper AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets (Lesci and Vlachos, 2024) published at the NAACL 2024 conference.
These artefacts can be reproduced using the code available at github.com/pietrolesci/anchoral.
The outputs/ folder includes the raw files created by the individual experiments.
The results/ folder contains the exported metrics and configurations that are used to complete the analysis and create the tables and… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/anchoral-paper-artefacts.claude-fable-5-claude-code
claude-fable-5 Agent Traces
It's worth noting that our team was working with Glint-Research to collect as much fable data as possible.
These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data).
For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.spreadsheet-arena-release
Spreadsheet Arena
A dataset of 555 pairwise human preference votes over LLM-generated spreadsheets, spanning 124 distinct user-submitted prompts and 17 models.
This is the public release accompanying the Spreadsheet Arena paper.
Contents
battles.csv
models.csv
outputs/<id>/
sheet.json
sheet.xlsx
<id> is a 16-char hex identifier (HMAC-SHA256 of an internal UUID under a… See the full description on the dataset page: https://huggingface.co/datasets/Longitude-Labs/spreadsheet-arena-release.ids-project-artifactsML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning.
The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering.
The dataset is maintained by with requests to the ArXiv API.
The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.DF-arrowGenImage_webp: Just convert all images of GenImage to webp format, just save disk space
GenImage_test: only contains raw GenImage test samples
GenImage: raw GenImage with all data
planck-2018-chains
Planck 2018 Cosmological-Parameter Chains
This dataset contains the Planck Public Release 3 cosmological-parameter
full grid, COM_CosmoParams_fullGrid_R3.01.zip. It contains the Markov
chains and their GetDist and CosmoMC companions for 336 combinations of
cosmological model and likelihood or external-data selection. The 1,296
chain roots become 5,184 Parquet tables containing 27,699,519 rows.
The tables retain the source's headerless, positional structure. Arrow
fields are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-2018-chains.CXM_Arena
Dataset Card for CXM Arena Benchmark Suite
Dataset Description
This dataset, "CXM Arena Benchmark Suite," is a comprehensive collection designed to evaluate various AI capabilities within the Customer Experience Management (CXM) domain. It consolidates five distinct tasks into a unified benchmark, enabling robust testing of models and pipelines in business contexts. The entire suite was synthetically generated using advanced large language models, primarily… See the full description on the dataset page: https://huggingface.co/datasets/sprinklr-huggingface/CXM_Arena.ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.agent-course-final-assignment
Agent Course Final Assignment - Unified Dataset
Author: Arte(r)m Sedov
GitHub: https://github.com/arterm-sedov/
Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment
Dataset Description
This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.Arabic_Aya
Dataset Card for : Arabic Aya (2A)
Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing
Dataset Sources & Infos
Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite.
Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp', 'acq' )… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/Arabic_Aya.filtering-pretraining-mix-arrow-formatarxivmath
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from ArXivMath used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
problem_type (list[string]): Problem type/category labels.
source (float64): arXiv… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/arxivmath.ledger-long-context-KPI-QA
LEDGER — Long-Context KPI Question Answering & Page Retrieval
This dataset is part of the LEDGER (Long-context Evaluation of Documents for
Grounded Extraction and Retrieval) benchmark.
It supports two of the three LEDGER tasks:
Page-level KPI retrieval — given a natural-language question about a financial
KPI and the corresponding annual report, retrieve the relevant page(s). Each row
includes TREC-style graded relevance judgments (qrels) over all candidate pages.… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-KPI-QA.
