datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolmino-mix-1124
DOLMino dataset mix for OLMo2 stage 2 annealing training.
Mixture of high-quality data used for the second stage of OLMo2 training.
Source Sizes
Name
Category
Tokens
Bytes (uncompressed)
Documents
License
DCLM
HQ Web Pages
752B
4.56TB
606M
CC-BY-4.0
Flan
HQ Web Pages
17.0B
98.2GB
57.3M
ODC-BY
Pes2o
STEM Papers
58.6B
413GB
38.8M
ODC-BY
Wiki
Encyclopedic
3.7B
16.2GB
6.17M
ODC-BY
StackExchange
CodeText
1.26B
7.72GB
2.48M
CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.all-defectsafrica-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.c4-en-html-with-training_metadata_allprosocial-dialog
Dataset Card for ProsocialDialog Dataset
Dataset Summary
ProsocialDialog is the first large-scale multi-turn English dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative… See the full description on the dataset page: https://huggingface.co/datasets/allenai/prosocial-dialog.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.vertebrate-v1-all
marin-dna/vertebrate-v1-all
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the all region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.hf-coding-tools-traces-all
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
31 sessions, one per (tool, model, effort, thinking) configuration
9,603 query → response turns total (≈19,206 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-all.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.gemma-2b-dictionary-embeddings-all-layers
Gemma-2B Dictionary Embeddings - All Layers
This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers.
Dataset Structure
metadata.json: Contains dataset metadata (model info, dimensions, word count)
embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26)
Usage
import pickle
from huggingface_hub import hf_hub_download
# Download a specific layer
layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.href_resultsDataDecide-eval-instances
DataDecide evaluation instances
This dataset contains data for individual evaluation instances
from the DataDecide project (publication forthcoming). It shows how
standard evaluation benchmarks can vary across many dimensions of
model design.
The dataset contains evaluations for a range of OLMo-style models
trained with:
25 different training data configurations
9 different sizes with parameter counts 4M, 20M, 60M, 90M, 150M, 300M, 750M, and 1B
3 initial random seeds
Multiple… See the full description on the dataset page: https://huggingface.co/datasets/allenai/DataDecide-eval-instances.Molmo2-TVQArobomme_1cuben_allcases4_phase4
robomme_1cuben_allcases4_phase4
Directly trainable LeRobot-format build of the exhaustive depth-4
VideoUnmaskSwap1CubeN dataset. A single red cube is hidden under one of three
fixed cups, the cups are shuffled zero to four times, and the robot must pick the
cup hiding the cube.
This repository includes the raw LeRobot parquet data, both precomputed camera
latents, the classifier-free empty text embedding, and deterministic train and
validation manifests. Unlike… See the full description on the dataset page: https://huggingface.co/datasets/alfayoung/robomme_1cuben_allcases4_phase4.lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private
Dataset Card for Evaluation run of allenai/llama-3-tulu-2-dpo-8b
Dataset automatically created during the evaluation run of model allenai/llama-3-tulu-2-dpo-8b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-allenai-llama-3-tulu-2-dpo-8b-private.bigym-all-tasks-3dgs-one-success
BiGym 全任务 3DGS 厨房壳 单成功轨迹集
发布状态:40/40 技术验证通过,视觉抽检通过;仓库保持,等待上游数据/壳资产再分发条款复核。
预览视频:previews/collection-preview.mp4
核心结果
官方 BiGym 任务:40/40
唯一任务:40
保存的 reward=1 episode:40
丢弃的 reward=0 候选:61(未进入成功 Parquet)
Transition / 每相机帧数:14,806 / 14,806
三相机 H.264 文件:120
相机:head 848×480、left wrist 640×480、right wrist 640×480,20 FPS
数据布局
reach/:3 个任务
long_horizon/:3 个任务
dishwasher/:10 个任务
tabletop/:24 个任务
task-manifest.csv:任务、demo UUID、seed、帧数、reward、动作哈希… See the full description on the dataset page: https://huggingface.co/datasets/eustance/bigym-all-tasks-3dgs-one-success.lm-eval-results-nlpguy-AlloyIngot-private
Dataset Card for Evaluation run of nlpguy/AlloyIngot
Dataset automatically created during the evaluation run of model nlpguy/AlloyIngot
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngot-private.lwm-spectro-alluserseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.lm-eval-results-Kukedlc-NeuralKuke-4-All-7b-private
Dataset Card for Evaluation run of Kukedlc/NeuralKuke-4-All-7b
Dataset automatically created during the evaluation run of model Kukedlc/NeuralKuke-4-All-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-NeuralKuke-4-All-7b-private.forbidden-backrooms-gemma-4-31B-it
Forbidden Backrooms: Gemma-4 31B Self-Chat
Self-chat transcripts and per-message embeddings for two role-inverted instances of Gemma-4-31B-it, comparing the official instruct checkpoint against an abliterated fine-tune of the same checkpoint. Both variants use identical int4 quantization served via Ollama, so quantization noise is not a confound between them.
The methodology follows Anthropic's Claude Opus 4 system card section on the "spiritual bliss attractor state." Leave two… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/forbidden-backrooms-gemma-4-31B-it.lm-eval-results-nlpguy-AlloyIngotNeo-private
Dataset Card for Evaluation run of nlpguy/AlloyIngotNeo
Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeo
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeo-private.kernelbench_with_promptsThis is a version of KernelBench where the prompts to produce the Triton and cuda kernel are explicitly saved in the JSON data files.
It only contains Level 1, 2, 3 kernels.
The prompt is the same as what is provided in the original KernelBench repo.
The dataset is prepared by Jiin Woo during her internship at AWS Annapurna Labs, the lab behind Trainium chips.
This dataset is part of an unreleased paper, and the paper will be updated in this README soon. If you use this dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/kernelbench_with_prompts.lm-eval-results-nlpguy-AlloyIngotNeoX-private
Dataset Card for Evaluation run of nlpguy/AlloyIngotNeoX
Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeoX
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeoX-private.pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.allenai__Llama-3.1-Tulu-3-70B-details
Dataset Card for Evaluation run of allenai/Llama-3.1-Tulu-3-70B
Dataset automatically created during the evaluation run of model allenai/Llama-3.1-Tulu-3-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__Llama-3.1-Tulu-3-70B-details.lm-eval-results-nlpguy-AlloyIngotNeoY-private
Dataset Card for Evaluation run of nlpguy/AlloyIngotNeoY
Dataset automatically created during the evaluation run of model nlpguy/AlloyIngotNeoY
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nlpguy-AlloyIngotNeoY-private.Sonicr_shortstories_24kFiltered and somewhat cleaned up scrape of posts from r/shortstories subreddit. Still has some reddit artifacts, but should be usable as is for training.
