CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolmino-mix-1124 DOLMino dataset mix for OLMo2 stage 2 annealing training. Mixture of high-quality data used for the second stage of OLMo2 training. Source Sizes Name Category Tokens Bytes (uncompressed) Documents License DCLM HQ Web Pages 752B 4.56TB 606M CC-BY-4.0 Flan HQ Web Pages 17.0B 98.2GB 57.3M ODC-BY Pes2o STEM Papers 58.6B 413GB 38.8M ODC-BY Wiki Encyclopedic 3.7B 16.2GB 6.17M ODC-BY StackExchange CodeText 1.26B 7.72GB 2.48M CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.tabulartext-generation100M<n<1B102 likes24k downloads11mo agoHugging Face02allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes20k downloads4y agoHugging Face03evalstate /all-defectstabularn<1K2 likes4.2k downloads5mo agoHugging Face04electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.7k downloads1mo agoHugging Face05bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes980 downloads3y agoHugging Face06allenai /prosocial-dialog Dataset Card for ProsocialDialog Dataset Dataset Summary ProsocialDialog is the first large-scale multi-turn English dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative… See the full description on the dataset page: https://huggingface.co/datasets/allenai/prosocial-dialog.tabulartext-classification100K<n<1M119 likes827 downloads4y agoHugging Face07FINAL-Bench /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.imagetext-generationn<1K26 likes811 downloads7mo agoHugging Face08marin-dna /vertebrate-v1-all marin-dna/vertebrate-v1-all Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the all region cohort with all species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.tabular100M<n<1B0 likes762 downloads2mo agoHugging Face09davidkling /hf-coding-tools-traces-all HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,603 query → response turns total (≈19,206 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-all.tabularn<1K0 likes648 downloads4mo agoHugging Face10youssef3146 /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.imagetext-generationn<1K0 likes482 downloads7mo agoHugging Face11LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes447 downloads1y agoHugging Face12allenai /href_resultstabularn<1K0 likes376 downloads1y agoHugging Face13allenai /DataDecide-eval-instances DataDecide evaluation instances This dataset contains data for individual evaluation instances from the DataDecide project (publication forthcoming). It shows how standard evaluation benchmarks can vary across many dimensions of model design. The dataset contains evaluations for a range of OLMo-style models trained with: 25 different training data configurations 9 different sizes with parameter counts 4M, 20M, 60M, 90M, 150M, 300M, 750M, and 1B 3 initial random seeds Multiple… See the full description on the dataset page: https://huggingface.co/datasets/allenai/DataDecide-eval-instances.tabular1K<n<10K2 likes330 downloads2y agoHugging Face14allenai /Molmo2-TVQAtabular100K<n<1M0 likes284 downloads7mo agoHugging Face15alfayoung /robomme_1cuben_allcases4_phase4 robomme_1cuben_allcases4_phase4 Directly trainable LeRobot-format build of the exhaustive depth-4 VideoUnmaskSwap1CubeN dataset. A single red cube is hidden under one of three fixed cups, the cups are shuffled zero to four times, and the robot must pick the cup hiding the cube. This repository includes the raw LeRobot parquet data, both precomputed camera latents, the classifier-free empty text embedding, and deterministic train and validation manifests. Unlike… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_allcases4_phase4.tabularroboticsn<1K0 likes264 downloads28d agoHugging Face16nyu-dice-lab /lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private Dataset Card for Evaluation run of allenai/llama-3-tulu-2-dpo-8b Dataset automatically created during the evaluation run of model allenai/llama-3-tulu-2-dpo-8b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private.tabular100K<n<1M0 likes263 downloads2y agoHugging Face17eustance /bigym-all-tasks-3dgs-one-success BiGym 全任务 3DGS 厨房壳 单成功轨迹集 发布状态:40/40 技术验证通过,视觉抽检通过;仓库保持,等待上游数据/壳资产再分发条款复核。 预览视频:previews/collection-preview.mp4 核心结果 官方 BiGym 任务:40/40 唯一任务:40 保存的 reward=1 episode:40 丢弃的 reward=0 候选:61(未进入成功 Parquet) Transition / 每相机帧数:14,806 / 14,806 三相机 H.264 文件:120 相机:head 848×480、left wrist 640×480、right wrist 640×480,20 FPS 数据布局 reach/:3 个任务 long_horizon/:3 个任务 dishwasher/:10 个任务 tabletop/:24 个任务 task-manifest.csv:任务、demo UUID、seed、帧数、reward、动作哈希… See the full description on the dataset page: https://huggingface.co/datasets/eustance/bigym-all-tasks-3dgs-one-success.tabularroboticsn<1K1 likes229 downloads2mo agoHugging Face18nyu-dice-lab /lm-eval-results-nlpguy-AlloyIngot-private Dataset Card for Evaluation run of nlpguy/AlloyIngot Dataset automatically created during the evaluation run of model nlpguy/AlloyIngot The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngot-private.tabular100K<n<1M0 likes177 downloads2y agoHugging Face19tomerraviv95 /lwm-spectro-allusertabularn<1K0 likes157 downloads3mo agoHugging Face20luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes130 downloads3mo agoHugging Face21nyu-dice-lab /lm-eval-results-Kukedlc-NeuralKuke-4-All-7b-private Dataset Card for Evaluation run of Kukedlc/NeuralKuke-4-All-7b Dataset automatically created during the evaluation run of model Kukedlc/NeuralKuke-4-All-7b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-NeuralKuke-4-All-7b-private.tabular100K<n<1M0 likes109 downloads2y agoHugging Face22alliedtoasters /forbidden-backrooms-gemma-4-31B-it Forbidden Backrooms: Gemma-4 31B Self-Chat Self-chat transcripts and per-message embeddings for two role-inverted instances of Gemma-4-31B-it, comparing the official instruct checkpoint against an abliterated fine-tune of the same checkpoint. Both variants use identical int4 quantization served via Ollama, so quantization noise is not a confound between them. The methodology follows Anthropic's Claude Opus 4 system card section on the "spiritual bliss attractor state." Leave two… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/forbidden-backrooms-gemma-4-31B-it.tabulartext-generation10K<n<100K0 likes99 downloads5mo agoHugging Face23nyu-dice-lab /lm-eval-results-nlpguy-AlloyIngotNeo-private Dataset Card for Evaluation run of nlpguy/AlloyIngotNeo Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeo The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeo-private.tabular100K<n<1M0 likes97 downloads2y agoHugging Face24allenanie /kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files. It only contains Level 1, 2, 3 kernels. The prompt is the same as what is provided in the original KernelBench repo. The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips. This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.tabularn<1K1 likes86 downloads1y agoHugging Face25nyu-dice-lab /lm-eval-results-nlpguy-AlloyIngotNeoX-private Dataset Card for Evaluation run of nlpguy/AlloyIngotNeoX Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeoX The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeoX-private.tabular100K<n<1M0 likes79 downloads2y agoHugging Face26novastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes79 downloads4mo agoHugging Face27open-llm-leaderboard /allenai__Llama-3.1-Tulu-3-70B-detailsgated Dataset Card for Evaluation run of allenai/Llama-3.1-Tulu-3-70B Dataset automatically created during the evaluation run of model allenai/Llama-3.1-Tulu-3-70B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__Llama-3.1-Tulu-3-70B-details.tabular10K<n<100K0 likes63 downloads2y agoHugging Face28nyu-dice-lab /lm-eval-results-nlpguy-AlloyIngotNeoY-private Dataset Card for Evaluation run of nlpguy/AlloyIngotNeoY Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeoY The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeoY-private.tabular100K<n<1M0 likes62 downloads2y agoHugging Face29Alltitude /Sonictabularn<1K0 likes61 downloads2y agoHugging Face30allura-org /r_shortstories_24kFiltered and somewhat cleaned up scrape of posts from r/shortstories subreddit. Still has some reddit artifacts, but should be usable as is for training. tabular10K<n<100K7 likes51 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.