datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking
Dataset Summary
WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information.
The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.EdgeBench
Overview
EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.oldi_seed
OLDI Seed Machine Translation Datacard
OLDI Seed is a machine translation dataset designed to be used to kick-start machine translation models for language directions which currently lack large-scale datasets.
Dataset Details
Dataset Description
OLDI Seed is a parallel corpus which consists of 6,193 sentences sampled from English Wikipedia and translated into 44 languages. It can be used to kick-start machine translation models for language… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/oldi_seed.GDPval-CN-Seed-Set
GDPval-CN Seed Set
中文详细说明 · English documentation · 样本说明
GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。
我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。
这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。
GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review.
数据概览
项目
内容
任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.unipic_seedream_4images
UniPic-Nano-4Images: A Multi-Image Composition Dataset
⚡ Quick Start
The image archive is split into multiple parts for easier downloading. To reconstruct and extract:
# Step 1: Concatenate split files into a single zip
cat nano-banana.part_* > nano-banana-4images.zip
# Step 2: Extract the images
unzip nano-banana-4images.zip
📖 Overview
UniPic-Nano-4Images is a high-quality multi-image composition dataset containing 48,805 samples designed for training… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_4images.agent-apprenticeship-seed-dataset
Agent Apprenticeship Seed Dataset
The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents.
As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.agent-apprenticeship-seed-dataset_v0.2
Agent Apprenticeship Seed Dataset v0.2
Real-world agent work experience, looped into collective learning.
The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset_v0.2.FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.sorrel-T-qwen3-8b-base-seed0-documentsunipic_seedream_6images
UniPic-Nano-6Images: A Complex Multi-Image Composition Dataset
⚡ Quick Start
The image archive is split into multiple parts for easier downloading. To reconstruct and extract:
# Step 1: Concatenate split files into a single zip
cat nano-banana.part_* > nano-banana-6images.zip
# Step 2: Extract the images
unzip nano-banana-6images.zip
📖 Overview
UniPic-Nano-6Images is a high-quality complex multi-image composition dataset containing 41,508 samples designed… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_6images.Primus-Seed
PRIMUS: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
🤗 Primus-Seed
Primus-Seed is a high-quality🚀 cybersecurity text dataset composed of data crawled from reputable sources such as MITRE, Wikipedia, and well-known cybersecurity company websites, as well as CTI manually collected by our threat experts.
Statistics
Category
Samples
Tokens
Avg.
Web Crawl / Official Dump
Cybersecurity Blogs/News
2,946
9,751,002
3… See the full description on the dataset page: https://huggingface.co/datasets/trendmicro-ailab/Primus-Seed.unipic_seedream_5images
UniPic-Nano-5Images: A Multi-Image Composition Dataset
⚡ Quick Start
The image archive is split into multiple parts for easier downloading. To reconstruct and extract:
# Step 1: Concatenate split files into a single zip
cat nano-banana.part_* > nano-banana-5images.zip
# Step 2: Extract the images
unzip nano-banana-5images.zip
📖 Overview
UniPic-Nano-5Images is a high-quality multi-image composition dataset containing 47,461 samples designed for training… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_5images.RL-seed-Decensor-Difficultysorrel-T-qwen3-1.7b-base-seed0-documentsopensec-seeds
OpenSec Seeds: Incident Response Scenarios for Agent Calibration
This dataset provides 220 taxonomy-stratified security incident scenarios for training and evaluating AI agents on incident response (IR) tasks. Each scenario includes entity definitions, attack kill chains, ground truth labels, and prompt injection payloads designed to test agent calibration under adversarial evidence.
Paper: OpenSec: Measuring Incident Response Agent Calibration Under Adversarial Evidence… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/opensec-seeds.sorrel-T-qwen3-14b-base-seed0-documentssorrel-T-olmo-2-32b-seed0-documentsSEED-Timeline-Annotations
Timeline Annotations for BONES-SEED Humanoid Motion Dataset
Dataset Description:
This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.ReSA
ReSA (Reasoned Safety Alignment)
Project Page | Paper
ReSA (Reasoned Safety Alignment) is an open-source synthetic safety-training dataset with 80K examples designed to enhance LLM robustness against jailbreak attacks through an "Answer-Then-Check" strategy. The dataset teaches models to first generate a summary of their intended answer, then critically evaluate its safety before providing a final response. This approach achieves superior safety performance while maintaining strong… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ReSA.sorrel-T-gemma-3-27b-pt-seed0-documentssorrel-T-mistral-small-24b-base-seed0-documentsmimicgen-square-d0-light-seed42-1000-opaque-rerender-lerobotseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.Seed
Data Instances
instance_id: (str) - A unique identifier for the instance
repo: (str) - The repository name including the owner
base_commit: (str) - The base commit hash where the reproduction is tested
date: (timestamp) - The date of the commit
project_name: (str) - The name of the project without owner
lang: (str) - The programming language of the repository
dockerfile: (str) - Dockerfile content
build_sh: (str) - Build script content
work_dir: (str) - Working directory path… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/Seed.SeedBench
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain.
GitHub page
Overview
SeedBench assesses LLMs across three core seed breeding stages:
Gene Information Retrieval
Gene Function and Regulation… See the full description on the dataset page: https://huggingface.co/datasets/yj12869741/SeedBench.DiscoX
DiscoX Translation Benchmark
DiscoX is a benchmark for the evaluation of LLMs on discourse- and expert-level translation tasks.
Dataset At A Glance
Languages: English ⇄ Chinese (100 English→Chinese tasks, 100 Chinese→English tasks)
Total samples: 200 discourse- and exprt-level translation items
Average passage length: ~1.7k characters (min 0.73k, max 3.04k)
Meta fields: primary & secondary domain labels, structured rubrics, prompt IDs,etc
Reference Rubrics: every task… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/DiscoX.spring-seed-catalog
Spring Seed Catalog
Cleared seed packets from the community garden intake.
Retained seed entries: 7
Featured seed: SEED-106 — lettuce / Green Wave
Harvest-year span: 2022-2025
Organic entries: 4
Crops (bean/carrot/lettuce/tomato): 2/3/1/1
sorrel-T2-qwen3-8b-base-seed0-documentsAInsteinBench
AInsteinBench
AInsteinBench is a benchmark for evaluating the capabilities of AI agents in solving scientific computing problems. It currently supports Einstein Toolkit and Multi-SWE-bench formats of coding questions.
📊 Dataset Overview
AInsteinBench provides 244 scientific computing tasks derived from multiple scientific repositories. These tasks have been verified on execution and also reviewed by corresponding domain experts to verify both software engineering and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/AInsteinBench.
