datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.SEED-Bench
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of SEED-Bench. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023seed,
title={Seed-bench: Benchmarking multimodal llms with generative comprehension}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench.regen_s3_200_seed30700
regen_s3_200_seed30700 — S3 (variety) staged-regen batch
200 successful tube-rack insertion episodes from Isaac Sim with a myCobot 320 M5, generated as
the S3 config of the four-way staged regen. S3 is S2 plus scene variety: five rack
configurations, two tube models, clutter and bore tubes.
Companions: S1,
S1b,
S2.
Contents
200 episodes, 200/200 success, 17,793 frames, 4.4 GB
69–138 steps per episode (mean 89.0), one 640×480 RGB PNG per step
Per episode:… See the full description on the dataset page: https://huggingface.co/datasets/corvinus-labs/regen_s3_200_seed30700.opencode_seed2.1_expert_skill_round_00_20260712regen_s1b_200_seed28700
regen_s1b_200_seed28700 — S1b (rotation + recovery) staged-regen batch
200 successful tube-rack insertion episodes from Isaac Sim with a myCobot 320 M5, generated as
the S1b config of the four-way staged regen (S1 / S1b / S2 / S3). S1b is S1 plus recovery
data and nothing else, so S1 vs S1b at equal N measures what the recovery demonstrations cost
in accuracy and buy in off-path behaviour.
Companion to regen_s1_200_seed28200,
which is the same configuration with no… See the full description on the dataset page: https://huggingface.co/datasets/corvinus-labs/regen_s1b_200_seed28700.RefOI-TLHFRefOI-TLHF: Token-Level Human Feedback for Referring Expressions
📃 Paper |🏠 Project Website
Overview
RefOI-TLHF is a companion dataset to RefOI, developed as part of the study "Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation."
This dataset focuses on token-level human feedback: for each referring expression—produced by either a human or a model—we annotate the minimal informative span that enables successful identification of the… See the full description on the dataset page: https://huggingface.co/datasets/Seed42Lab/RefOI-TLHF.SEED-Bench-2
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of SEED-Bench-2. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@article{li2023seed2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/SEED-Bench-2.VisualWebInstruct-Seed
Introduction
This is the seed dataset we used to conduct Google Search.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
SimAct
SimAct
SimAct is a synthetic action-variation image dataset generated from the MSCOCO dataset. For each COCO source image, the dataset adds 4 or 5 generated images depicting different actions or action-like scene changes.
Data Fields
Each row contains:
source_image_id: COCO source image id
image: generated image
source_image: original COCO source image
type: generated or original image (for this dataset, all generated)
action: short action description
description:… See the full description on the dataset page: https://huggingface.co/datasets/Seed42Lab/SimAct.RefBlocksdxl_images_easy_prompts-artists-seed1sainfoin-seed-datasettext-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.MathCanvas_gemini3flash_seed42_idx_900_1050ByteCameraDepth
ByteCameraDepth Dataset
Paper | Project Page | Code
ByteCameraDepth is a multi-camera depth estimation dataset containing synchronized depth, color, and auxiliary data captured from various 3D cameras. The dataset provides comprehensive depth sensing from multiple cameras in various in-door scenarios, making it ideal for developing and evaluating depth estimation algorithms.
Dataset Overview
Purpose: Multi-camera depth estimation research and benchmarking
Total… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ByteCameraDepth.laion2b_seed
Dataset Card for "laion2b_seed"
This dataset is a subset of laion2B-en-aesthetic, with SEED v1 tokens.
SEED-Bench-2-Plusfrom https://huggingface.co/datasets/AILab-CVC/SEED-Bench-2-plus
SEED-Bench-2-Plus Card
Benchmark details
Benchmark type: SEED-Bench-2-Plus is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs). It consists of 2.3K multiple-choice questions with precise human annotations, spanning three broad categories: Charts, Maps, and Webs, each of which covers a wide spectrum of text-rich scenarios in the real world.
Benchmark date: SEED-Bench-2-Plus was collected in April 2024.… See the full description on the dataset page: https://huggingface.co/datasets/doolayer/SEED-Bench-2-Plus.Env-seed-media
Env Seed Media
Binary seed assets (images, video, audio, fonts, brand marks) for the
sandbox environments used by the DecodingTrust agent platform and the
forgingground environment suite.
These files are separated out of the application repositories so the
env source stays lightweight and text-only. Each environment fetches its
media from here at build/seed time and places it back under the paths shown
below.
Layout
The tree mirrors each environment's own… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/Env-seed-media.K-SEED
K-SEED
We introduce K-SEED, a Korean adaptation of the SEED-Bench [1] designed for evaluating vision-language models.
By translating the first 20 percent of the test subset of SEED-Bench into Korean, and carefully reviewing its naturalness through human inspection, we developed a novel robust evaluation benchmark specifically for Korean language.
K-SEED consists of questions across 12 evaluation dimensions, such as scene understanding, instance identity, and instance attribute… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/K-SEED.Seedream-3_t2i_human_preference
Rapidata Seedream 3 Preference
This T2I dataset contains over ~400'000 human responses from over ~30'000 individual annotators, collected in less than 7h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating OpenAI 4o (version from 26.3.2025) across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Seedream-3_t2i_human_preference.prof_images_blip__SD_v1.4_random_seeds
Dataset Card for "prof_images_blip__SD_v1.4_random_seeds"
More Information needed
seed_sacha_inchi
Seed Sacha Inchi Dataset
Deskripsi Dataset
Dataset ini berisi citra biji sacha inchi yang diklasifikasikan ke dalam dua kelas, yaitu bagus dan rusak. Dataset ini disusun untuk mendukung penelitian dan pengembangan model Computer Vision, khususnya pada tugas klasifikasi kualitas visual biji sacha inchi.
Kelas Dataset
Dataset terdiri dari dua kelas:
Label
Deskripsi
bagus
Citra biji sacha inchi dengan kondisi visual baik
rusak
Citra biji… See the full description on the dataset page: https://huggingface.co/datasets/masdenpur/seed_sacha_inchi.regen_s2_200_seed29700
regen_s2_200_seed29700 — S2 (appearance) staged-regen batch
200 successful tube-rack insertion episodes from Isaac Sim with a myCobot 320 M5, generated as
the S2 config of the four-way staged regen. S2 is S1b plus full appearance randomization
and nothing else, so S1b vs S2 at equal N measures what appearance DR costs and buys.
Companions: S1 (rotation
only), S1b (+ recovery).
Contents
200 episodes, 200/200 success, 17,939 frames, 4.8 GB
63–159 steps per episode… See the full description on the dataset page: https://huggingface.co/datasets/corvinus-labs/regen_s2_200_seed29700.SeeDo
License
The SeeDo dataset is licensed under the Creative Commons
Attribution 4.0 International (CC BY 4.0) License.
This dataset is self-curated in the IROS 2025 paper: VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model (https://arxiv.org/abs/2410.08792)
If you use this dataset in your research, please cite:
@inproceedings{wang2025vlm,
title={Vlm see, robot do: Human demo video to robot action plan via vision language model},
author={Wang… See the full description on the dataset page: https://huggingface.co/datasets/ai4ce/SeeDo.regen_s1_plus400_seed32000SEED-Bench-2-plus
SEED-Bench-2-Plus Card
Benchmark details
Benchmark type:
SEED-Bench-2-Plus is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs).
It consists of 2.3K multiple-choice questions with precise human annotations, spanning three broad categories: Charts, Maps,
and Webs, each of which covers a wide spectrum of text-rich scenarios in the real world.
Benchmark date:
SEED-Bench-2-Plus was collected in April 2024.
Paper or resources for more information:… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Bench-2-plus.image-to-video-human-preference-seedance-1-pro
Rapidata Video Generation Hailuo-02 v Marey Human Preference
In this dataset, ~6k human responses from ~2k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 5 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-to-video-human-preference-seedance-1-pro.regen_s1b_plus400_seed33000ScienceOlympiad
ScienceOlympiad: Challenging AI with Olympiad-Level Multimodal Science Problems
Dataset Description
The ScienceOlympiad dataset is a meticulously curated benchmark designed to test the limits of current AI models in scientific reasoning. It comprises elite, competition-level problems in physics and chemistry. Addressing the need for more diverse and realistic challenges, ScienceOlympiad introduces multimodal integration as a key dimension. Unlike purely text-based… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ScienceOlympiad.BM-6M-Demo
Dataset Card for ByteMorph-6M-Demo
The task of editing images to reflect non-rigid motions, such as changes in camera viewpoint, object deformation, human articulation, or complex interactions, represents a significant yet underexplored frontier in computer vision. Current methodologies and datasets often concentrate on static imagery or rigid transformations, thus limiting their applicability to expressive edits involving dynamic movement. To bridge this gap, we present… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo.
