datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt-perfect
Scoring popular datasets with "Self-Alignment with Instruction Backtranslation" prompt
35 datasets scored (>6B tokens)
Scoring Models used
gpt-3.5-turbo-16k
gpt-3.5-turbo-1106
gpt-3.5-turbo-0125
All datasets have 2 additional columns
score - Response from the model including CoT (if provided)
extracted_score - Extracted score from the score column as int
Datasets Scored by Prompt (Needs to be updated)… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/prompt-perfect.SWE-Hero-openhands-trajectories
SWE-Hero Trajectories: Execution-based Fine-tuning for Software Engineering Agents
Data Overview
SWE-Hero Trajectories is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 34k agent
trajectories collected using the OpenHands framework. The trajectories
were synthesized using Qwen3-Coder-480B-A35B-Instruct, specifically curated for supervised fine-tuning (SFT),
aiming to improve model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SWE-Hero-openhands-trajectories.heronstegoattack-advbench50
StegoAttack AdvBench-50
Steganographic jailbreak data generated using the StegoAttack pipeline from the paper "Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks" (Geng et al., 2025).
For experiment results and analysis, see experiment.md.
What is StegoAttack?
StegoAttack is a jailbreak method that uses steganography to hide harmful queries inside benign-looking text. It embeds each word of a harmful query at a fixed position (e.g. the 2nd… See the full description on the dataset page: https://huggingface.co/datasets/heron-ai-security/stegoattack-advbench50.swe-mt-combined-coderforge-hero-lego-nex-swezero
fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (7)
togethercomputer__CoderForge-Preview
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.hero_run_4_math_codeherorun1_code-test_50K_150K
Dataset card for herorun1_code-test_50K_150K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are tasked with implementing a softmax layer for a neural network using CUDA and C++. The softmax layer is a common component in neural network architectures and is used to normalize the output of a network to a probability distribution over multiple classes.\n\nYour task is to implement the following CUDA kernel for… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_50K_150K.Qwen3-32B_hero_run_4_code_32k-sharegptswe-mt-combined-hero-lego-nex-swezero
fan-shu/swe-mt-combined-hero-lego-nex-swezero
Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded
in order and concatenated into a single config so one training epoch visits every
trajectory exactly once (no interleave / no oversampling).
Built from fan-shu/swe-instruct-trajectories-empty-think-inserted.
Source subsets (6)
nvidia__SWE-Zero-openhands-trajectories
nex-agi__agent-sft
nvidia__SWE-Hero-openhands-trajectories… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-hero-lego-nex-swezero.isic2017_task3Matter-0.1
Matter 0.1
Curated top quality records from 35 other datasets. Extracted from prompt-perfect
This is just a consolidation of all the score 5s. Fine-tuning models with various subsets and combinations to create a best performing v1 dataset
~1.4B Tokens, ~2.5M records
Dataset has been deduped, decontaminated with bagel script from Jon Durbin
Download using the below command to avoid unecessary files
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/Matter-0.1.Japanese-Heron-BenchThis dataset is a clarified version of the image, context, and question set included in the Japanese-Heron-Bench for the construction of the Japanese evaluation benchmark suite.
The original dataset refers to turing-motors/Japanese-Heron-Bench.
Link to the original dataset🔗: https://huggingface.co/datasets/turing-motors/Japanese-Heron-Bench
@misc{inoue2024heronbench,
title={Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese},
author={Yuichi Inoue and Kento… See the full description on the dataset page: https://huggingface.co/datasets/Silviase/Japanese-Heron-Bench.Matter-0.2-alphadetails_0-hero__Matter-0.2-7B
Dataset Card for Evaluation run of 0-hero/Matter-0.2-7B
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.2-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_0-hero__Matter-0.2-7B.Japanese-Heron-Bench
Japanese-Heron-Bench
Dataset Description
Japanese-Heron-Bench is a benchmark for evaluating Japanese VLMs (Vision-Language Models). We collected 21 images related to Japan. We then set up three categories for each image: Conversation, Detail, and Complex, and prepared one or two questions for each category. The final evaluation dataset consists of 102 questions. Furthermore, each image is assigned one of seven subcategories: anime, art, culture, food, landscape, landmark… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japanese-Heron-Bench.details_0-hero__Matter-0.2-32B
Dataset Card for Evaluation run of 0-hero/Matter-0.2-32B
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.2-32B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_0-hero__Matter-0.2-32B.place_bottleOIG-small-chip2
Dataset Card for "OIG-small-chip2"
OIG-small-chip2 dataset from https://laion.ai/blog/oig-dataset/
Original Dataset - https://github.com/LAION-AI/Open-Instruction-Generalist
openthoughts3_herorun_ckpt06500_eval_27e9
mlfoundations-dev/openthoughts3_herorun_ckpt06500_eval_27e9
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 59.95% ± 0.79%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
59.30%
303
511
2
62.82%
321
511
3
60.08%
307
511
4
58.51%
299
511
5
61.45%
314
511
6
57.53%
294
511
HEROCS_Datasetstocks-HEROMOTOCO-1D-candlesHerosHEROS is a dataset used to compare the sentence cosine similarity among sentences with high lexical overlapping but differ in their semantics.
Please refer to the paper, "Revealing the Blind Spot of Sentence Encoder Evaluation by HEROS" for more details of how the dataset is constructed and the comparison of different sentence encoders.
The dataset heros.tsv consists of 6 columns: Original, Synonym, Antonym, Negation, Random, Typo, Negation.
The first column, Original are the sentences from… See the full description on the dataset page: https://huggingface.co/datasets/dcml0714/Heros.btme-news-heroeshard-hat-heroes
Hard Hat Heroes: Construction Safety Detection
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
4,000
Labeled training data
test
1,000
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
image
Image
image_id
string
width
int64
height
int64
objects.bbox… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hard-hat-heroes.herondetails_0-hero__Matter-0.1-Slim-7B-A
Dataset Card for Evaluation run of 0-hero/Matter-0.1-Slim-7B-A
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.1-Slim-7B-A on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_0-hero__Matter-0.1-Slim-7B-A.lm-eval-results-nbeerbower-HeroBophades-2x7B-private
Dataset Card for Evaluation run of nbeerbower/HeroBophades-2x7B
Dataset automatically created during the evaluation run of model nbeerbower/HeroBophades-2x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-HeroBophades-2x7B-private.herorun1_code-test_150K_250K
Dataset card for herorun1_code-test_150K_250K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "**A Recursive Riddle** \nImaginary scenario: I'm part of an experimental team at a tech lab where our latest project involves constructing a recursively defined program that reveals its own architecture. The mission is to build a function that not only discloses how many layers of functions exist but also specifies the code… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_150K_250K.details_0-hero__Matter-0.1-Slim-7B-B
Dataset Card for Evaluation run of 0-hero/Matter-0.1-Slim-7B-B
Dataset automatically created during the evaluation run of model 0-hero/Matter-0.1-Slim-7B-B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_0-hero__Matter-0.1-Slim-7B-B.herorun1_code-test_250K_350K
Dataset card for herorun1_code-test_250K_350K
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"problem": "You are tasked with creating a contract in Solidity that includes various utility functions related to token calculations. The contract should include functions for retrieving and setting decimals for tokens, calculating destination and source amounts based on token rates, retrieving token balances, and performing various… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/herorun1_code-test_250K_350K.
