datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XDof-TshirtFolding-20hours-normalizedmc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
oscar2301-ja-filter-ja-normal
Dataset Card for "oscar2301-ja-filter-ja-normal"
More Information needed
Depth-Normal-Images-617Kcv24-tr-128-normalizedDepth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedcv24-cy-128-normalizedNormalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
textures-color-normal-1k
textures-color-normal-1k
Dataset Summary
The textures-color-normal-1k dataset is an image dataset of 1000+ color and normal map textures in 512x512 resolution.
The dataset was created for use in image to image tasks.
It contains a combination of CC0 procedural and photoscanned PBR materials from ambientCG.
Dataset Structure
Data Instances
Each data point contains a 512x512 color texture and the corresponding 512x512 normal map.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/dream-textures/textures-color-normal-1k.cv24-ur-128-normalizedcv24-uk-128-normalizedcv24-pt-128-normalizedcv24-sw-128-normalizedcv24-de-128-normalizedrouter-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.coco2017_caption_normalcv24-sk-128-normalizedstack-v2-python-normal-onlylichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedcv24-ca-128-normalizedmerged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 800,
"total_frames": 133851,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task.epsilon-normalized
Dataset Card for "epsilon-normalized"
More Information needed
gspc-normalized
GSPC normalised — every bank in one schema
The one schema to read first. 518 rows that flatten several GSPC banks into a
single shape: source (the bank repository the row came from), axis, category, anchor, prompt,
expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently
dropped. If you want to reuse the banks without learning each one's native layout, start here.
The live board is the authority
GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.igbo_tts_normalizedv1v1d_docmatix_1k_normal_v6Skill-normalisation-ESCO-graded
skill-normalisation-esco-graded
Graded-relevance annotations for surface skill terms (ESCO alt-labels) from
ESCO v1.1.0 skill-normalisation pairs
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
50
_id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.v1v1d_docmatix_1k_normal_v5salesforce-xlam-finetune-normal
