datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEED-Bench
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of SEED-Bench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023seed,
title={Seed-bench: Benchmarking multimodal llms with generative comprehension}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench.Code-Contests-Plus
CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases
Introduction
CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions.
Highlights
High… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus.WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking
Dataset Summary
WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information.
The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.THEMol
THEMol: Torsion, Hessian, Energy of Molecules
Dataset Summary
THEMol is an open-source collection of quantum mechanical properties tailored for organic molecules. It provides large-scale density functional theory (DFT) data for exploring intramolecular potential energy surfaces, including optimized geometries, structural relaxation trajectories, torsion scans, constrained torsion relaxation trajectories, Hessian matrices, and MBIS-derived atomic properties.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/THEMol.obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.BeyondAIME
BeyondAIME: Advancing Math Reasoning Evaluation Beyond High School Olympiads
Dataset Description
BeyondAIME is a curated test set designed to benchmark advanced mathematical reasoning. Its creation was guided by the following core principles to ensure a fair and challenging evaluation:
High Difficulty: Problems are sourced from high-school and university mathematics competitions, with a difficulty level greater than or equal to that of AIME Problems #11-15.… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BeyondAIME.TerminalWorld-Seeds-Clean
TerminalWorld Seeds, oracle-validated
Current TW curation selection (20260911-v4)
The current TW selection retains 405 tasks, excluding any original TW task
excluded by either the preserved local quality curation or the upstream v3 training
selection. The TW-only training JSONL is the Hub default config.
For mixed training, use the 858-row JSONL:405 TW plus all453 upstream
admitted TMAX tasks unchanged (mixed-curated config).
This release preserves the approved… See the full description on the dataset page: https://huggingface.co/datasets/andylizf/TerminalWorld-Seeds-Clean.mga-fineweb-edu
Massive Genre-Audience Augment Fineweb-Edu Corpus
This dataset is a synthetic pretraining corpus described in paper Reformulation for Pretraining Data Augmentation.
Overview of synthesis framework. Our method expands the original corpus through a two-stage synthesis process.
Each document is reformulated to 5 new documents, achieving 3.9× token number expansion while maintaining diversity through massive (genre, audience) pairs.
We build MGACorpus based on SmolLM Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/mga-fineweb-edu.opencode_seed2.1_expert_skill_round_00_20260712TerminalWorld-Seeds
TerminalWorld Seeds, packaged in the RST release layout
1,530 validated terminal tasks from EuniAI/TerminalWorld, repackaged in the release layout of Zhongzhi1228/Recursive-Task-Synthesis (RST, arXiv:2608.05466).
RST bootstrapped its recursive synthesis from 639 seeds sampled out of TerminalWorld, but released only the synthesized rounds. This dataset is the seed-level superset in the same format, so a synthesis pipeline can start from round 0 with the same loaders that read the… See the full description on the dataset page: https://huggingface.co/datasets/andylizf/TerminalWorld-Seeds.RefOI-TLHFRefOI-TLHF: Token-Level Human Feedback for Referring Expressions
📃 Paper |🏠 Project Website
Overview
RefOI-TLHF is a companion dataset to RefOI, developed as part of the study "Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation."
This dataset focuses on token-level human feedback: for each referring expression—produced by either a human or a model—we annotate the minimal informative span that enables successful identification of the… See the full description on the dataset page: https://huggingface.co/datasets/Seed42Lab/RefOI-TLHF.SEED-Bench-2
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of SEED-Bench-2. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023seed2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench-2.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.VisualWebInstruct-Seed
Introduction
This is the seed dataset we used to conduct Google Search.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
SimAct
SimAct
SimAct is a synthetic action-variation image dataset generated from the MSCOCO dataset. For each COCO source image, the dataset adds 4 or 5 generated images depicting different actions or action-like scene changes.
Data Fields
Each row contains:
source_image_id: COCO source image id
image: generated image
source_image: original COCO source image
type: generated or original image (for this dataset, all generated)
action: short action description
description:… See the full description on the dataset page: https://huggingface.co/datasets/Seed42Lab/SimAct.seed-tts-eval-50-arrowoldi_seed
OLDI Seed Machine Translation Datacard
OLDI Seed is a machine translation dataset designed to be used to kick-start machine translation models for language directions which currently lack large-scale datasets.
Dataset Details
Dataset Description
OLDI Seed is a parallel corpus which consists of 6,193 sentences sampled from English Wikipedia and translated into 44 languages. It can be used to kick-start machine translation models for language… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/oldi_seed.ruler-300-seed42
Frozen RULER 300, seed 42
This dataset freezes the exact RULER inputs used by the
short-long-pretraining native evaluation suite.
Repository: bicycleman15/ruler-300-seed42
Rows: 6,300
Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq
Context lengths: 1024, 2048, 4096
Samples per task/length: 300
Seed: 42
Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d
Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.sdxl_images_easy_prompts-artists-seed1sainfoin-seed-datasetGDPval-CN-Seed-Set
GDPval-CN Seed Set
中文详细说明 · English documentation · 样本说明
GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。
我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。
这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。
GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review.
数据概览
项目
内容
任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.opencode_seed2.1_expert_without_reproduce_round_00text-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.BiointelligenceAgent01-SeedDataset
Biointelligence Agent Worlds Seed Dataset 0.1
Configurations
Configuration
Contents
Rows
targets
Public-real and explicitly synthetic targets
1000
entities
Normalized world entities
14516
world_snapshots
Observed and simulated temporal states
356014
source_events
Retrieved/discovered source records and normalized claims
33074
relationships
Evidence-linked target and entity relationships
205800
agent_runs
Observable structured Codex run results… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/BiointelligenceAgent01-SeedDataset.multitask_german_examples_32kSWE-Smith-Seeds-Clean
SWE-Smith Seeds, agent-verified
1,552 of SWE-smith's 59,136 instances, repackaged as terminal tasks and kept only where every claim about them was executed and held: the bug is present, the reference fix earns the grader's reward, the repository's own suite still passes, and a coding agent solved the task from its instruction alone in a sandbox that had neither the fix nor the tests nor the network. Every row carries the verdict and the conditions it was taken under; nothing… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/SWE-Smith-Seeds-Clean.ByteCameraDepth
ByteCameraDepth Dataset
Paper | Project Page | Code
ByteCameraDepth is a multi-camera depth estimation dataset containing synchronized depth, color, and auxiliary data captured from various 3D cameras. The dataset provides comprehensive depth sensing from multiple cameras in various in-door scenarios, making it ideal for developing and evaluating depth estimation algorithms.
Dataset Overview
Purpose: Multi-camera depth estimation research and benchmarking
Total… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ByteCameraDepth.SEEDBench2
Dataset Card for Dataset Name
The SEEDBench2 evaluation dataset hosted by VLMEval (authorized by the author).
Dataset Details
Language(s) (NLP): English
License: Apache 2.0
Repository: https://github.com/AILab-CVC/SEED-Bench
Paper [optional]: https://arxiv.org/abs/2311.17092
Citation
@misc{li2023seedbench2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models},
author={Bohao Li and Yuying Ge and Yixiao Ge and Guangzhi Wang and Rui… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/SEEDBench2.
