datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt-perfect
Scoring popular datasets with "Self-Alignment with Instruction Backtranslation" prompt
35 datasets scored (>6B tokens)
Scoring Models used
gpt-3.5-turbo-16k
gpt-3.5-turbo-1106
gpt-3.5-turbo-0125
All datasets have 2 additional columns
score - Response from the model including CoT (if provided)
extracted_score - Extracted score from the score column as int
Datasets Scored by Prompt (Needs to be updated)… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/prompt-perfect.SWE-Hero-openhands-trajectories
SWE-Hero Trajectories: Execution-based Fine-tuning for Software Engineering Agents
Data Overview
SWE-Hero Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 34k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SWE-Hero-openhands-trajectories.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.hero_run_4_math_codeherorun1_code-test_50K_150K
Dataset card for herorun1_code-test_50K_150K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are tasked with implementing a softmax layer for a neural network using CUDA and C++. The softmax layer is a common component in neural network architectures and is used to normalize the output of a network to a probability distribution over multiple classes.\n\nYour task is to implement the following CUDA kernel for… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_50K_150K.swe-mt-combined-hero-lego-nex-swezero
fan-shu/swe-mt-combined-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (6)
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft
nvidia__SWE-Hero-openhands-trajectories… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-hero-lego-nex-swezero.Matter-0.1
Matter 0.1
Curated top quality records from 35 other datasets. Extracted from prompt-perfect
This is just a consolidation of all the score 5s. Fine-tuning models with various subsets and combinations to create a best performing v1 dataset
~1.4B Tokens, ~2.5M records
Dataset has been deduped, decontaminated with bagel script from Jon Durbin
Download using the below command to avoid unecessary files
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/Matter-0.1.Japanese-Heron-BenchThis dataset is a clarified version of the image, context, and question set included in the Japanese-Heron-Bench for the construction of the Japanese evaluation benchmark suite.
The original dataset refers to turing-motors/Japanese-Heron-Bench.
Link to the original dataset🔗: https://huggingface.co/datasets/turing-motors/Japanese-Heron-Bench
@misc{inoue2024heronbench,
title={Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese},
author={Yuichi Inoue and Kento… See the full description on the dataset page: https://huggingface.co/datasets/Silviase/Japanese-Heron-Bench.Matter-0.2-alphaOIG-small-chip2
Dataset Card for "OIG-small-chip2"
OIG-small-chip2 dataset from https://laion.ai/blog/oig-dataset/
Original Dataset - https://github.com/LAION-AI/Open-Instruction-Generalist
openthoughts3_herorun_ckpt06500_eval_27e9
mlfoundations-dev/openthoughts3_herorun_ckpt06500_eval_27e9
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 59.95% ± 0.79%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
59.30%
303
511
2
62.82%
321
511
3
60.08%
307
511
4
58.51%
299
511
5
61.45%
314
511
6
57.53%
294
511
HerosHEROS is a dataset used to compare the sentence cosine similarity among sentences with high lexical overlapping but differ in their semantics.
Please refer to the paper, "Revealing the Blind Spot of Sentence Encoder Evaluation by HEROS" for more details of how the dataset is constructed and the comparison of different sentence encoders.
The dataset heros.tsv consists of 6 columns: Original, Synonym, Antonym, Negation, Random, Typo, Negation.
The first column, Original are the sentences from… See the full description on the dataset page: https://huggingface.co/datasets/dcml0714/Heros.hard-hat-heroes
Hard Hat Heroes: Construction Safety Detection
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
4,000
Labeled training data
test
1,000
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
image
Image
image_id
string
width
int64
height
int64
objects.bbox… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hard-hat-heroes.lm-eval-results-nbeerbower-HeroBophades-2x7B-private
Dataset Card for Evaluation run of nbeerbower/HeroBophades-2x7B
Dataset automatically created during the evaluation run of model nbeerbower/HeroBophades-2x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-HeroBophades-2x7B-private.herorun1_code-test_150K_250K
Dataset card for herorun1_code-test_150K_250K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "**A Recursive Riddle** \nImaginary scenario: I'm part of an experimental team at a tech lab where our latest project involves constructing a recursively defined program that reveals its own architecture. The mission is to build a function that not only discloses how many layers of functions exist but also specifies the code… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_150K_250K.herorun1_code-test_250K_350K
Dataset card for herorun1_code-test_250K_350K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are tasked with creating a contract in Solidity that includes various utility functions related to token calculations. The contract should include functions for retrieving and setting decimals for tokens, calculating destination and source amounts based on token rates, retrieving token balances, and performing various… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_250K_350K.SWE-Hero-openhands-trajectories
SWE-Hero Trajectories: Execution-based Fine-tuning for Software Engineering Agents
Data Overview
SWE-Hero Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 34k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/SWE-Hero-openhands-trajectories.prompt-perfect-dpo
DPO Version of Prompt Perfect
Update
02-22-2024
Noticed a correlation with the rejected_pair generation prompt (or scoring) where length of response (level of detail) is almost proportional to quality.
Testing new prompts for a re-run where is quality is not directly proportional to length of response directly
This might result in models that generate long responses
All datasets have 4 additional columns
accepted_pair - Original… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/prompt-perfect-dpo.hero_run_4_codeherorun1_code-test_350K_450K
Dataset card for herorun1_code-test_350K_450K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are tasked with simulating the behavior of a simple text-based user interface using a series of commands. The interface consists of a single line of text, and the commands are executed in the order they are given. Each command is in the format `./send_command.py <action> <arguments>`, where `<action>` is the action to be… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_350K_450K.lm-eval-results-nbeerbower-HeroBophades-3x7B-private
Dataset Card for Evaluation run of nbeerbower/HeroBophades-3x7B
Dataset automatically created during the evaluation run of model nbeerbower/HeroBophades-3x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-HeroBophades-3x7B-private.Matter-0.1-Slim-D0-hero__Matter-0.2-7B-DPO-details
Dataset Card for Evaluation run of 0-hero/Matter-0.2-7B-DPO
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.2-7B-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/0-hero__Matter-0.2-7B-DPO-details.openthoughts3_herorun_ckpt04500_eval_636d
mlfoundations-dev/openthoughts3_herorun_ckpt04500_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
63.0
91.3
89.8
0.3
64.9
47.3
53.8
24.0
25.3
AIME24
Average Accuracy: 63.00% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
56.67%
17
30
2
60.00%
18
30
3
63.33%
19
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_herorun_ckpt04500_eval_636d.herorun1_code-test_30K_50K
Dataset card for herorun1_code-test_30K_50K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are working on a machine learning project and need to implement a method for updating model parameters and random effect realizations. The code snippet provided is a part of a Python class that handles these operations. Your task is to create a function that updates the model and random effect realization attributes based on… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_30K_50K.herorun3_codeherorun1_code_0-25000
Dataset card for herorun1_code_0-25000
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "For today's challenge, you are tasked with developing a program or function that transforms a string by reversing the order of vowels while maintaining the positions of consonants and non-alphabetic characters. The transformation should ensure that the output string has its original structure, with vowels appearing in the reverse… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code_0-25000.herorun3_mathMatter-0.1-Slim-BSubset B of Matter-0.1
Datasets have been deduped, decontaminated with the bagel script from Jon Durbin
distilabel-math-preference-dpo
