CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ByteDance-Seed /Code-Contests-Plus CodeContests+: A Competitive Programming Dataset with High-Quality Test Cases Introduction CodeContests+ is a competitive programming problem dataset built upon CodeContests. It includes 11,690 competitive programming problems, along with corresponding high-quality test cases, test case generators, test case validators, output checkers, and more than 13 million correct and incorrect solutions. Highlights High… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Code-Contests-Plus.tabularother10K<n<100K69 likes13k downloads11mo agoHugging Face02ByteDance-Seed /THEMol THEMol: Torsion, Hessian, Energy of Molecules Dataset Summary THEMol is an open-source collection of quantum mechanical properties tailored for organic molecules. It provides large-scale density functional theory (DFT) data for exploring intramolecular potential energy surfaces, including optimized geometries, structural relaxation trajectories, torsion scans, constrained torsion relaxation trajectories, Hessian matrices, and MBIS-derived atomic properties. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/THEMol.tabular10M<n<100M7 likes10k downloads4mo agoHugging Face03andylizf /TerminalWorld-Seeds TerminalWorld Seeds, packaged in the RST release layout 1,530 validated terminal tasks from EuniAI/TerminalWorld, repackaged in the release layout of Zhongzhi1228/Recursive-Task-Synthesis (RST, arXiv:2608.05466). RST bootstrapped its recursive synthesis from 639 seeds sampled out of TerminalWorld, but released only the synthesized rounds. This dataset is the seed-level superset in the same format, so a synthesis pipeline can start from round 0 with the same loaders that read the… See the full description on the dataset page: https://huggingface.co/datasets/andylizf/TerminalWorld-Seeds.tabular1K<n<10K0 likes1.6k downloads1mo agoHugging Face04olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes1.2k downloads4y agoHugging Face05AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.2k downloads2y agoHugging Face06EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.tabular10K<n<100K0 likes957 downloads27d agoHugging Face07EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.tabular10K<n<100K0 likes910 downloads27d agoHugging Face08EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.tabular10K<n<100K0 likes893 downloads27d agoHugging Face09bicycleman15 /ruler-300-seed42 Frozen RULER 300, seed 42 This dataset freezes the exact RULER inputs used by the short-long-pretraining native evaluation suite. Repository: bicycleman15/ruler-300-seed42 Rows: 6,300 Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq Context lengths: 1024, 2048, 4096 Samples per task/length: 300 Seed: 42 Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.tabularquestion-answering1K<n<10K0 likes875 downloads1mo agoHugging Face10EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.tabular10K<n<100K0 likes872 downloads27d agoHugging Face11Rapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes708 downloads1y agoHugging Face12rayrren /BiointelligenceAgent01-SeedDataset Biointelligence Agent Worlds Seed Dataset 0.1 Configurations Configuration Contents Rows targets Public-real and explicitly synthetic targets 1000 entities Normalized world entities 14516 world_snapshots Observed and simulated temporal states 356014 source_events Retrieved/discovered source records and normalized claims 33074 relationships Evidence-linked target and entity relationships 205800 agent_runs Observable structured Codex run results… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/BiointelligenceAgent01-SeedDataset.tabularother100K<n<1M0 likes584 downloads12d agoHugging Face13seedboxai /multitask_german_examples_32ktabular100K<n<1M15 likes567 downloads3y agoHugging Face14Fzz1 /SWE-Smith-Seeds-Clean SWE-Smith Seeds, agent-verified 1,552 of SWE-smith's 59,136 instances, repackaged as terminal tasks and kept only where every claim about them was executed and held: the bug is present, the reference fix earns the grader's reward, the repository's own suite still passes, and a coding agent solved the task from its instruction alone in a sandbox that had neither the fix nor the tests nor the network. Every row carries the verdict and the conditions it was taken under; nothing… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/SWE-Smith-Seeds-Clean.tabular1K<n<10K0 likes535 downloads1mo agoHugging Face15EleutherAI /LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005 Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005 This repository contains 100 fully retrained language models, not just scores. Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents) of the 4,656-document WikiText-2 training set from EleutherAI/bergson-wikitext-2-4656-chunks, following the recipe of Bae et al. 2024, Training Data Attribution via Approximate Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.tabular10K<n<100K0 likes469 downloads27d agoHugging Face16liangyuch /laion2b_seed Dataset Card for "laion2b_seed" This dataset is a subset of laion2B-en-aesthetic, with SEED v1 tokens. image100M<n<1B1 likes436 downloads3y agoHugging Face17rayrren /agent-apprenticeship-seed-dataset Agent Apprenticeship Seed Dataset The living ecosystem where AI agents run automated workflow loops on any task, improve through execution, and turn each run into reusable work experience + data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where real-world tasks generate reusable learning signals and complex workflows advance through agent loops that turn execution into shared… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset.tabular1K<n<10K0 likes425 downloads3mo agoHugging Face18rayrren /agent-apprenticeship-seed-dataset_v0.2 Agent Apprenticeship Seed Dataset v0.2 Real-world agent work experience, looped into collective learning. The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents. As agents move into long-horizon, economically valuable work, Agent Apprenticeship creates the open infrastructure where… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/agent-apprenticeship-seed-dataset_v0.2.tabular10K<n<100K0 likes401 downloads3mo agoHugging Face19JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes312 downloads2mo agoHugging Face20yjernite /prof_report__SD_v2_random_seeds__multi__24 Dataset Card for "prof_report__SD_v2_random_seeds__multi__24" More Information needed tabular1K<n<10K0 likes306 downloads3y agoHugging Face21hazyresearch /OT_8K_seed_all_responsestabular100K<n<1M0 likes282 downloads11mo agoHugging Face22ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.tabulartext-generation10K<n<100K0 likes256 downloads3mo agoHugging Face23Jarrodbarnes /opensec-seeds OpenSec Seeds: Incident Response Scenarios for Agent Calibration This dataset provides 220 taxonomy-stratified security incident scenarios for training and evaluating AI agents on incident response (IR) tasks. Each scenario includes entity definitions, attack kill chains, ground truth labels, and prompt injection payloads designed to test agent calibration under adversarial evidence. Paper: OpenSec: Measuring Incident Response Agent Calibration Under Adversarial Evidence… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/opensec-seeds.tabularreinforcement-learningn<1K1 likes240 downloads7mo agoHugging Face24ai-safety-institute /reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2) GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking. Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.tabulartext-generation10K<n<100K0 likes224 downloads3mo agoHugging Face25liangyuch /laion2B-en-aesthetic-seed Dataset Card for "laion2B-en-aesthetic-seed" More Information needed image1M<n<10M3 likes222 downloads3y agoHugging Face26Lilambd /world-seeds World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2 Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join. One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.tabular1K<n<10K0 likes215 downloads2h agoHugging Face27Lyrasilas /eval_ep100_seedNone_circle_big_car_less_zoom_40000_defaultThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "racecar", "total_episodes": 20, "total_frames": 6599, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Lyrasilas/eval_ep100_seedNone_circle_big_car_less_zoom_40000_default.tabularrobotics1K<n<10K0 likes209 downloads7mo agoHugging Face28synpre /dclm_seed_5b_tanishqtabular1M<n<10M0 likes192 downloads2y agoHugging Face29Lyrasilas /eval_ep1000_seedNone_default_car_guessed_10000_SFT_circle_bigThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "racecar", "total_episodes": 20, "total_frames": 18779, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Lyrasilas/eval_ep1000_seedNone_default_car_guessed_10000_SFT_circle_big.tabularrobotics10K<n<100K0 likes190 downloads7mo agoHugging Face30Lyrasilas /eval_ep500_seed1_default_center_20000_defaultThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "racecar", "total_episodes": 20, "total_frames": 12019, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Lyrasilas/eval_ep500_seed1_default_center_20000_default.tabularrobotics10K<n<100K0 likes184 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.