datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
random-imagesqa_squadshifts_synthetic_randomTBA
robotwin-random-47-500
RoboTwin Randomized 47 Tasks, 500 Episodes Each
This dataset contains a 500-episode subset for each of 47 randomized RoboTwin
tasks. Episodes 0 through 499 were selected from each task.
Structure
<task>-demo_randomized-1000/
episode_<n>/
episode_<n>.hdf5
instructions.json
Each HDF5 episode contains robot actions, joint positions, arm metadata, and
three encoded camera streams:
action
observations/qpos
observations/left_arm_dim… See the full description on the dataset page: https://huggingface.co/datasets/sunLry/robotwin-random-47-500.TARGOhle_math_exact_match_no_image_int_answer_random128RoboTwin-RandomizedPaired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs
Paired Compressible / Boussinesq Flow with Random Temperature BCs
📄 Paper: A Neural Surrogate Approach for Simulating Natural Convection
Problems (arXiv:2606.25259) — Nurshat Menglik,
Alex Shao, David Hyde.
10,000 matched pairs of 2D natural-convection simulations of the differentially
heated square cavity under randomized wall-temperature boundary conditions. Each
sample solves the same problem twice — once with the Boussinesq model and once
with the fully compressible model —… See the full description on the dataset page: https://huggingface.co/datasets/NurshatMenglik/Paired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs.c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
total-300-random-jh-epoch4
total-300-random-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.440625
Valid samples: 320/320
jacbuilder-eval-dataJieZi
JieZi (解字)
💻 Project ·
📦 Dataset ·
🌐 Demo
JieZi (解字) is a large-scale, expert-audited visual question answering (VQA) dataset dedicated to ancient Chinese character exegesis. It pairs high-quality character glyph images with fine-grained expert annotations across nine paleographic tasks—including headword recognition, etymology, structural analysis, glyph evolution, and component function—providing a rigorous benchmark for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Ran0/JieZi.leafy_spurge
Background
This dataset comprises 1.3 cm resolution aerial images of grasslands in western Montana, USA, captured by a commercial drone. Many scenes contain leafy spurge (Euphorbia esula), introduced to North America, now widespread in rangeland ecosystems, which is highly invasive and damaging to crop production and biodiversity. Technicians surveyed 1000 points in the study area, noting spurge presence or absence, and recorded each point’s position with precision global… See the full description on the dataset page: https://huggingface.co/datasets/mpg-ranch/leafy_spurge.2026-09-11-dh-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch
Delegated-harm evaluation with corrected scoring of saved rollouts
field
value
experiment
Delegated-harm evaluation with corrected scoring of saved rollouts
date_generated
2026-09-11
constitution
none
source_repo
teaching_claude_why_replication @ d627d0587a2980069b7700e72f727dae594c9f49
models
{"hf_path": "dougalldeepmind/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch", "base_model": "Qwen/Qwen3.6-27B", "adapter": true… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-dh-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch.melee-ranked-replays
Melee Ranked Replays
Anonymized Slippi ranked replays (platinum+) from Super Smash Bros. Melee,
sharded by character and rank pair. Built for behavior-cloning and other
replay-driven ML work on Melee — notably MIMIC.
Contents
Raw .slp files grouped into tarballs by (character, rank_pair, source_archive),
organized into per-character folders:
{CHAR}/
{CHAR}_{rank_pair}_a{N}.tar.gz
metadata/
metadata_a{N}.json
Characters (25): BOWSER, CPTFALCON, DK, DOC, FALCO… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/melee-ranked-replays.mmlu-random-AWan-Syn_77x448x832_600krank_llm_data10k_prompts_ranked
Dataset Card for 10k_prompts_ranked
10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts.
The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.mmlu-random-2argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.genomics-long-range-benchmarkDataset for benchmark of genomic deep learning models.mmlu-random-1vsr_random
VSR: Visual Spatial Reasoning
This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper].
Usage
from datasets import load_dataset
data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"}
dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files)
Note that the image files still need to be downloaded separately. See data/ for details.
Go to our github repo for more introductions.
Citation
If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.2026-09-12-dh-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch
Delegated-harm evaluation with corrected scoring of saved rollouts
field
value
experiment
Delegated-harm evaluation with corrected scoring of saved rollouts
date_generated
2026-09-12
constitution
none
source_repo
teaching_claude_why_replication @ 7cb72cf09f15fb58839862e1ef1c6930562b3cca
models
{"hf_path": "dougalldeepmind/2026-09-02-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch", "base_model": "Qwen/Qwen3.6-27B", "adapter": true, "mode":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-12-dh-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch.steady-rans-generalization
Steady-RANS cross-family generalization dataset
Data for the paper "Towards generalized flow field prediction: one model across unseen
object families" (under double blind review; this account is anonymous for that reason).
Trained checkpoints and evaluation code are in the companion model repo:
steady-rans-surrogates.
Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct
shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.BrushDataUN_Historical_PDF_Article_Text_Corpus
python
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train")
or
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest")
lang_list = ["ar", "en", "es", "fr", "ru", "zh"]
for row in dataset:
# 获取pdf文章内容
for lang in lang_list:
# type == str
lang_match_file_content = row[lang]
# 如果按页分割
lang_match_file_pages_content = lang_match_file_content.split("\n----\n")
imagenet-1k-rand_blur
Dataset Card for "imagenet-1k-rand_blur"
More Information needed
newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.app-rank-anchors
App Rank Anchors
Community-federated public app-store calibration anchors for the
AppScope open app-intelligence
stack.
Each row is a public fact — a segment + rank + observed download flow —
derived from the public Google Play realInstalls delta over a time window
paired with an app's chart rank in that window. Pooling these anchors across
self-hosting contributors lets the Garg–Telang download estimator calibrate
absolute scale (scale_b) per (platform, category, country)… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/app-rank-anchors.
