datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XDof-TshirtFolding-20hours-normalizedmmu-norm-legacy-north
Legacy Survey DR9 North image cutouts — L1 (release v1)
This L1 repository contains 2,191,927 objects matched across the release, in 726 shards (about 700 GB). Each object has 152×152-pixel cutouts at 0.262″ per pixel in three bands.
Band names: DR9 North imaging uses BASS g/r and MzLS z. This repository uses bass-g, bass-r, and mzls-z; the original des-g/des-r/des-z tokens are kept in native_band_tokens.
Schema
image struct: band (3), flux (3×152×152… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-legacy-north.OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedDepth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
Depth-Normal-Images-617KOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedlichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.mmu-norm-hsc
HSC PDR3 Deep/UltraDeep image cutouts — L1 (release v1)
This L1 repository contains 55,346 objects matched across the release, in 79 shards (about 57 GB). Each object has 160×160-pixel cutouts at 0.168″ per pixel in five bands (hsc-g/r/i/z/y), with per-pixel inverse variance and a mask.
Schema
image struct: band (5), flux (ADU at AB zeropoint 27), ivar, mask (true = CLEAN — verified empirically), psf_fwhm, scale. Plus cmodel magnitudes/errors, extendedness… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-hsc.NormasTCU
NormasTCU
Overview
NormasTCU iis a dataset for Legal Information Retrieval (LIR) in Brazilian Portuguese composed of normative documents from the Brazilian Federal Court of Accounts (Tribunal de Contas da União - TCU), along with queries and human-annotated relevance judgments.
The dataset includes:
14,469 legal documents (normative acts);
46 queries;
812 judge query-document pairs derived from 3,048 human annotations with 3-level graded relevance.… See the full description on the dataset page: https://huggingface.co/datasets/LeandroRibeiro/NormasTCU.mmu-norm-sne
Supernova & transient photometry, 11 catalogs — L1 (release v1)
This L1 repository contains 627 events matched across the release from YSE DR1, SNLS, Swift, Foundation, PS1, DES-Y3, CSP, CfA3, CfA4, CfA stripped-envelope, and CfA SN II. Each catalog has its own Parquet file, with a survey column added. The measurements retain their native values.
Schema
Per event: lightcurve struct with time, band, and either flux/flux_err (YSE, SNLS, Swift, Foundation, PS1, DES)… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-sne.E_normal_over70_add
Dataset Card for "E_normal_over70_add"
More Information needed
VFD_normalize_9_v1meld-open-normalized
MELD Open (Normalized)
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.norman-perturb
Norman 2019 CRISPRa K562 perturbation atlas
CRISPRa (gene activation) in K562 chronic myeloid leukemia cells. Single-cell expression in log-normalized counts.
Generated 2026-06-08 as one of three companion atlases (Norman, Replogle, VCC).
File schema (each config / single-config repo)
File
Shape
Description
pseudobulks.h5ad
(50, n_genes)
50 control pseudobulks (15 cells each, log-normalized means). Cell-type-specific baseline.
coexpression.h5ad… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-lynn/norman-perturb.grabette-tactile-normal-4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-4.voxpopuli_asr_norm_curatormmu-norm-tess
TESS-SPOC light curves — reacquired L1 (release v1)
This L1 repository contains 31,058 per-sector light curves for 5,482 matched TIC objects. The data were reacquired from the TESS-SPOC HLSP at STScI (archive.stsci.edu/hlsps/tess-spoc), with every available sector stored separately alongside the native QUALITY bits and float64 times.
Schema (one row per star-sector)
Column
Meaning
tic
TIC identifier
sector
TESS sector — sectors are separate rows… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-tess.normal_ball_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 44089,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/satvikahuja/normal_ball_full.Stuffed_Animal_V4.1_3cam_Normal_bboxes
Stuffed_Animal_V4.1_3cam_Normal
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
grabette-tactile-normal-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-2.grabette-tactile-normal-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal-3.pusht_96_norm2_remap
pusht_96_norm2
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2_remap.normalcamera_HKWSThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 16843,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zhaoraning/normalcamera_HKWS.mmu-norm-desi
DESI EDR (SV3) spectra — L1 (release v1)
This L1 repository contains 1,060,507 spectra matched across the release, in 21 shards (about 80 GB).
Schema
spectrum struct per row: lambda (vacuum, barycentric, observer frame, Å; fixed 7,781-sample grid), flux (×10⁻¹⁷ erg s⁻¹ cm⁻² Å⁻¹), ivar, lsf_sigma (approximate line-spread σ per pixel), mask (native; nonzero = bad — verified), valid (mask==0 AND ivar>0 AND finite(flux)). Plus object_id (TARGETID), ra, dec, and source… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-desi.E_normal_over70
Dataset Card for "E_normal_over70"
More Information needed
pusht_96_norm4
pusht_96_norm4
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
200
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4.mmu-norm-sdss
SDSS spectra — L1 (release v1)
This L1 repository contains 271,966 spectra matched across the release, in 11 shards (about 15 GB).
Schema
spectrum struct per row: lambda (vacuum, heliocentric, observer frame, Å; variable length ≤ ~4,700), flux (×10⁻¹⁷ erg s⁻¹ cm⁻² Å⁻¹), ivar, lsf_sigma, mask (native; nonzero = bad — verified), valid. Plus object_id, positions, metadata.
Padding: the upstream arrays used lambda = −1 with zero inverse variance for padding. Those… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-sdss.grabette-tactile-normalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 17,
"total_frames": 6800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:17"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CarolinePascal/grabette-tactile-normal.scientific_lay_summarisation-plos-norm
scientific_lay_summarisation - PLOS - normalized
This dataset is a modified version of tomasg25/scientific_lay_summarization and contains scientific lay summaries that have been preprocessed with this code. The preprocessing includes fixing punctuation and whitespace problems, and calculating the token length of each text sample using a tokenizer from the T5 model.
Original dataset details:
Repository: https://github.com/TGoldsack1/Corpora_for_Lay_Summarisation
Paper: Making… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/scientific_lay_summarisation-plos-norm.CartonPickNPlace2Target-normalized
