datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.quokka_ckptsEnglish-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.InternData-fractal20220817_dataswallow-code-v2
SwallowCode-v2
Resources
📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881.
🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning.
💻 What is it?
SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline.
However, it had two significant limitations:
(1) it was distributed under the Llama 3.3 Community License, and
(2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.stackv2_edu_filtered
Stack V2 Edu
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.hplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.data_2
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.osworld_v2_tasks
OSWorld V2 Task Classes
This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2.
The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task.
Download from the public repository root with:
uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4125
Action score: 0.4265625
Valid samples: 320/320
qwen35-4b-reeval2
qwen35-4b-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4265625
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.40546875
Action score: 0.475
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4703125
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.39921875
Action score: 0.44375
Valid samples: 320/320
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38359375
Action score: 0.4703125
Valid samples: 320/320
imagenet_hard_review_data_r2math-contests-2026
Math Contests 2026 (🔗 notadib/math-contests-2026)
197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data.
Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
spreadsheet-bench-v2-modified
SpreadsheetBench V2 Modified: Multi-Document QA
1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents.
An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.Japanese_NicoNico_Douga_Movie_Meta_Data_2016JMedBench
Maintainers
Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp
Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp
If you find any error in this benchmark or want to contribute to this benchmark, please feel free to contact us.
Introduction
This is a dataset collection of JMedBench, which is a benchmark for evaluating Japanese biomedical large language models (LLMs).
Details can be found in this paper. We also provide an evaluation framework, med-eval… See the full description on the dataset page: https://huggingface.co/datasets/Coldog2333/JMedBench.StreamAudio-2M
StreamAudio-2M
Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a
stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips
are organised into six task subsets.
Subsets
Subset
Rows
Description
Stream_Audio_Understanding
90,738
Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA
Real_time_ASR
28,109
Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page: https://huggingface.co/datasets/zhifeixie/StreamAudio-2M.asl-now-fingerspelling
ASLNow!
ASLNow! is a web app designed to make learning ASL fingerspelling easy and fun! You can try it live at asl-now.vercel.app.
Demo: https://www.youtube.com/watch?v=Wi5tAxVasq8
Dataset
This dataset, used to train the fingerspelling model is licensed under the MIT License. It will be updated frequently as more data is collected.
The dataset is collected from multiple participants told to sign ASL letters into a camera and detecting hand landmarks
using
the Mediapipe… See the full description on the dataset page: https://huggingface.co/datasets/sid220/asl-now-fingerspelling.xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.
