datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.CulturaY-ja-askllm-v1
CulturaY-ja-askllm-v1
多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.human_genomeSLT-Task2-Post-ASR-Speaker-Tagging
Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization)
Description
This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system.
Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging.
SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.text-code-galeras-code-generation-from-docstring-3k-dedupedGenDS
[CVPR-2025] GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration
Dataset Card for GenDS dataset
The GenDS dataset is a large dataset to boost the generalization of image restoration models. It is a combination of existing image restoration datasets and
diffusion-generated degraded samples from GenDeg.
Usage
The dataset is fairly large at ~360GB. We recommend having at least 800GB of free space. To download the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sudarshan2002/GenDS.reward-projection-goal-generalisation-vlmGEN
Human, AI-Generated, and AI-Edited Text: Stylometric Corpus
📄 Paper: https://arxiv.org/pdf/2608.27855
💻 Code: https://github.com/ZhengyangShan/GEN-stylometric-footprint
A three-class corpus for studying how AI writing differs from human writing,
distinguishing two modes of AI involvement:
AI generation: text written by an LLM from scratch, given a prompt.
AI editing: human text revised by an LLM (grammar, tone, paraphrase, etc.).
Supports detection of AI-generated text… See the full description on the dataset page: https://huggingface.co/datasets/szyszy/GEN.OpenMathInstruct-1-1.8m-ja-askllm-v1
OpenMathInstruct-1-1.8m-ja-askllm-v1
データセット kunishou/OpenMathInstruct-1-1.8m-ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/OpenMathInstruct-1-1.8m-ja-askllm-v1.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.Nemotron-RLHF-GenRM-v1
Dataset Description:
This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking.
The dataset is composed of:
Preference data focused on diverse domains
A synthetic safety blend
The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.genomes-brassicales-balanced-v1More info: https://github.com/songlab-cal/gpn
mammalian-genomes-cdsseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.genai-incidents
GenAI & Agentic AI Security Incidents
13,060 real-world and research incidents involving generative-AI and agentic-AI
systems — prompt injection, jailbreaks, data exfiltration, deepfakes, model and
supply-chain compromise, agent hijacking, and AI-enabled harms — cross-mapped to six
taxonomies. Dataset version 2.10.0.
Every applicable incident is tagged with four core taxonomies:
OWASP Top 10 for LLM Applications (2026) — LLM01–LLM10
OWASP Agentic Top 10 (ASI) — ASI01–ASI10
NIST… See the full description on the dataset page: https://huggingface.co/datasets/emmanuelgjr/genai-incidents.bat_genomeAngiosperm_16_genomesGene-Embedding-Hub-contributions-stagingSkill2-Bench
Skill²-Bench
Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning".
Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.GenixerForShikra-DatasetsPaper: Genixer: Empowering Multimodal Large Language Model as a Powerful Data Generator (ECCV 2024)
Arxiv: https://arxiv.org/abs/2312.06731
Description: syn_lcs_filtered60.jsonl and syn_sbu_filtered60.jsonl are two synthetic datasets produced by our Genixer_S model for advancing grounding-based multimodal understanding.
CodeForcesWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeForces.repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces
OpticalDNA reproduction — Codex agent trace
This dataset contains the raw Codex JSONL session trace for the ICML 2026
reproduction of Rethinking Genomic Modeling Through Optical Character
Recognition.
Published Trackio logbook
Paper page
Challenge instructions
Agent Trace Viewer announcement
The JSONL is uploaded directly from the matching ~/.codex/sessions entry, as
recommended by the Agent Trace Viewer. It captures the reproduction work,
Hugging Face Jobs audit, poster… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/repro-rethinking-genomic-modeling-through-optical-character-recognition-agent-traces.GeneralAgentBench
GeneralAgentBench
GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review.
Each task provides a natural-language instruction plus a list of verification
checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.CodeContests_trainWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests_train.zelk12__MT2-Gen7-gemma-2-9B-details
Dataset Card for Evaluation run of zelk12/MT2-Gen7-gemma-2-9B
Dataset automatically created during the evaluation run of model zelk12/MT2-Gen7-gemma-2-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zelk12__MT2-Gen7-gemma-2-9B-details.LiveBench-ReasonFluxWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/LiveBench-ReasonFlux.CodeContestsWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/CodeContests.MBPP-ReasonFluxWe use Stdio input/output format here. For example, for the task to calculate the sum of a list, the input and output are in the following format:
input = "5\n1 2 3 4 5\n"
output = "15"
CodeContests and CodeForces are using this format, however, MBPP and part of LiveCodeBench are using functional input/output format, such like
assert sum_function([1, 2, 3, 4, 5]) == 15
In this project, we have converted the the functional format to the Stdio format to achieve consistency.
Paper | Code… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/MBPP-ReasonFlux.GenRA-human-eval
GenRA — Human-Evaluation Study Data
This repository contains the stimuli, protocol, and anonymized ratings of the
blind pairwise human-evaluation study reported in the GenRA paper, in which
GenRA is compared against the closest state-of-the-art baseline
Puppeteer [Song et al., 2025].
It is released to support the reproducibility and scrutiny of the reported
results.
Study design
We sample 40 rigged assets from Articulation-XL-2.0
[Song et al., 2025;
HF dataset]… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GenRA-human-eval.
