CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ByteDance-Seed /WideSearch WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent's ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information. The challenge in these tasks lies not in… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/WideSearch.textn<1K45 likes12k downloads1y agoHugging Face02ByteDance-Seed /EdgeBench Overview EdgeBench is a benchmark of 134 real-world tasks for evaluating how autonomous AI agents learn from real-world environments. Instead of measuring one-shot performance, EdgeBench places agents in executable task environments with realistic, multi-level feedback and lets them iterate for 12+ hours per task — tracking the full trajectory of improvement, not just the final score. We publicly release 51 tasks… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/EdgeBench.texttext-generationn<1K84 likes7.5k downloads2mo agoHugging Face03AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.2k downloads2y agoHugging Face04openlanguagedata /oldi_seed OLDI Seed Machine Translation Datacard OLDI Seed is a machine translation dataset designed to be used to kick-start machine translation models for language directions which currently lack large-scale datasets. Dataset Details Dataset Description OLDI Seed is a parallel corpus which consists of 6,193 sentences sampled from English Wikipedia and translated into 44 languages. It can be used to kick-start machine translation models for language… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/oldi_seed.texttext-generation100K<n<1M11 likes965 downloads3mo agoHugging Face05human-intelligence-ai /GDPval-CN-Seed-Set GDPval-CN Seed Set 中文详细说明 · English documentation · 样本说明 GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。 我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。 这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。 GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review. 数据概览 项目 内容 任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.documentothern<1K1 likes790 downloads2mo agoHugging Face06Skywork /unipic_seedream_4images UniPic-Nano-4Images: A Multi-Image Composition Dataset ⚡ Quick Start The image archive is split into multiple parts for easier downloading. To reconstruct and extract: # Step 1: Concatenate split files into a single zip cat nano-banana.part_* > nano-banana-4images.zip # Step 2: Extract the images unzip nano-banana-4images.zip 📖 Overview UniPic-Nano-4Images is a high-quality multi-image composition dataset containing 48,805 samples designed for training… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_4images.textimage-to-image10K<n<100K5 likes456 downloads8mo agoHugging Face07rayrren /agent-apprenticeship-seed-dataset Agent Apprenticeship Seed Dataset The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.tabular1K<n<10K0 likes425 downloads3mo agoHugging Face08rayrren /agent-apprenticeship-seed-dataset_v0.2 Agent Apprenticeship Seed Dataset v0.2 Real-world agent work experience, looped into collective learning. The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset_v0.2.tabular10K<n<100K0 likes401 downloads3mo agoHugging Face09m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes379 downloads2y agoHugging Face10joshycodes /sorrel-T-qwen3-8b-base-seed0-documentstext100K<n<1M0 likes330 downloads7d agoHugging Face11Skywork /unipic_seedream_6images UniPic-Nano-6Images: A Complex Multi-Image Composition Dataset ⚡ Quick Start The image archive is split into multiple parts for easier downloading. To reconstruct and extract: # Step 1: Concatenate split files into a single zip cat nano-banana.part_* > nano-banana-6images.zip # Step 2: Extract the images unzip nano-banana-6images.zip 📖 Overview UniPic-Nano-6Images is a high-quality complex multi-image composition dataset containing 41,508 samples designed… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_6images.textimage-to-image10K<n<100K3 likes327 downloads8mo agoHugging Face12trendmicro-ailab /Primus-Seedgated PRIMUS: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training 🤗 Primus-Seed Primus-Seed is a high-quality🚀 cybersecurity text dataset composed of data crawled from reputable sources such as MITRE, Wikipedia, and well-known cybersecurity company websites, as well as CTI manually collected by our threat experts. Statistics Category Samples Tokens Avg. Web Crawl / Official Dump Cybersecurity Blogs/News 2,946 9,751,002 3… See the full description on the dataset page: https://huggingface.co/datasets/trendmicro-ailab/Primus-Seed.texttext-generation100K<n<1M27 likes321 downloads1y agoHugging Face13Skywork /unipic_seedream_5images UniPic-Nano-5Images: A Multi-Image Composition Dataset ⚡ Quick Start The image archive is split into multiple parts for easier downloading. To reconstruct and extract: # Step 1: Concatenate split files into a single zip cat nano-banana.part_* > nano-banana-5images.zip # Step 2: Extract the images unzip nano-banana-5images.zip 📖 Overview UniPic-Nano-5Images is a high-quality multi-image composition dataset containing 47,461 samples designed for training… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/unipic_seedream_5images.textimage-to-image10K<n<100K3 likes314 downloads8mo agoHugging Face14NewEden /RL-seed-Decensor-Difficultytext10K<n<100K1 likes301 downloads20d agoHugging Face15joshycodes /sorrel-T-qwen3-1.7b-base-seed0-documentstext100K<n<1M0 likes282 downloads7d agoHugging Face16Jarrodbarnes /opensec-seeds OpenSec Seeds: Incident Response Scenarios for Agent Calibration This dataset provides 220 taxonomy-stratified security incident scenarios for training and evaluating AI agents on incident response (IR) tasks. Each scenario includes entity definitions, attack kill chains, ground truth labels, and prompt injection payloads designed to test agent calibration under adversarial evidence. Paper: OpenSec: Measuring Incident Response Agent Calibration Under Adversarial Evidence… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/opensec-seeds.tabularreinforcement-learningn<1K1 likes240 downloads7mo agoHugging Face17joshycodes /sorrel-T-qwen3-14b-base-seed0-documentstext100K<n<1M0 likes212 downloads7d agoHugging Face18joshycodes /sorrel-T-olmo-2-32b-seed0-documentstext100K<n<1M0 likes193 downloads7d agoHugging Face19nvidia /SEED-Timeline-Annotations Timeline Annotations for BONES-SEED Humanoid Motion Dataset Dataset Description: This dataset provides additional text description annotations from the BONES-SEED humanoid motion dataset. For each motion, this dataset provides an overview text description of the entire motion at a high level, along with a “timeline” of annotated segments within the motion. Each segment generally contains a single atomic action and is defined by a start time, end time, and text… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SEED-Timeline-Annotations.text100K<n<1M7 likes191 downloads6mo agoHugging Face20ByteDance-Seed /ReSA ReSA (Reasoned Safety Alignment) Project Page | Paper ReSA (Reasoned Safety Alignment) is an open-source synthetic safety-training dataset with 80K examples designed to enhance LLM robustness against jailbreak attacks through an "Answer-Then-Check" strategy. The dataset teaches models to first generate a summary of their intended answer, then critically evaluate its safety before providing a final response. This approach achieves superior safety performance while maintaining strong… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/ReSA.texttext-generation10K<n<100K6 likes179 downloads7mo agoHugging Face21joshycodes /sorrel-T-gemma-3-27b-pt-seed0-documentstext100K<n<1M0 likes179 downloads7d agoHugging Face22joshycodes /sorrel-T-mistral-small-24b-base-seed0-documentstext100K<n<1M0 likes160 downloads7d agoHugging Face23insagur /mimicgen-square-d0-light-seed42-1000-opaque-rerender-lerobottabularn<1K0 likes157 downloads7mo agoHugging Face24luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes142 downloads3mo agoHugging Face25SEC-bench /Seed Data Instances instance_id: (str) - A unique identifier for the instance repo: (str) - The repository name including the owner base_commit: (str) - The base commit hash where the reproduction is tested date: (timestamp) - The date of the commit project_name: (str) - The name of the project without owner lang: (str) - The programming language of the repository dockerfile: (str) - Dockerfile content build_sh: (str) - Build script content work_dir: (str) - Working directory path… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/Seed.textn<1K0 likes131 downloads1y agoHugging Face26yj12869741 /SeedBench SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. GitHub page Overview SeedBench assesses LLMs across three core seed breeding stages: Gene Information Retrieval Gene Function and Regulation… See the full description on the dataset page: https://huggingface.co/datasets/yj12869741/SeedBench.textquestion-answering1K<n<10K1 likes130 downloads1y agoHugging Face27ByteDance-Seed /DiscoX DiscoX Translation Benchmark DiscoX is a benchmark for the evaluation of LLMs on discourse- and expert-level translation tasks. Dataset At A Glance Languages: English ⇄ Chinese (100 English→Chinese tasks, 100 Chinese→English tasks) Total samples: 200 discourse- and exprt-level translation items Average passage length: ~1.7k characters (min 0.73k, max 3.04k) Meta fields: primary & secondary domain labels, structured rubrics, prompt IDs,etc Reference Rubrics: every task… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/DiscoX.texttranslationn<1K5 likes114 downloads10mo agoHugging Face28SOTAagi2030 /spring-seed-catalog Spring Seed Catalog Cleared seed packets from the community garden intake. Retained seed entries: 7 Featured seed: SEED-106 — lettuce / Green Wave Harvest-year span: 2022-2025 Organic entries: 4 Crops (bean/carrot/lettuce/tomato): 2/3/1/1 tabularn<1K0 likes113 downloads4d agoHugging Face29joshycodes /sorrel-T2-qwen3-8b-base-seed0-documentstext10K<n<100K0 likes103 downloads7d agoHugging Face30ByteDance-Seed /AInsteinBench AInsteinBench AInsteinBench is a benchmark for evaluating the capabilities of AI agents in solving scientific computing problems. It currently supports Einstein Toolkit and Multi-SWE-bench formats of coding questions. 📊 Dataset Overview AInsteinBench provides 244 scientific computing tasks derived from multiple scientific repositories. These tasks have been verified on execution and also reviewed by corresponding domain experts to verify both software engineering and… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/AInsteinBench.tabulartext-generation1K<n<10K4 likes101 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.