datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.T2I-ImageNet-Normalmc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
router-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.oscar2301-ja-filter-ja-normal
Dataset Card for "oscar2301-ja-filter-ja-normal"
More Information needed
Depth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
Normalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
coco2017_caption_normalDepth-Normal-Images-617Kcv24-tr-128-normalizedcv24-cy-128-normalizedOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedcv24-ur-128-normalizedcv24-pt-128-normalizedcv24-uk-128-normalizedcv24-sw-128-normalizedcv24-de-128-normalizedcv24-sk-128-normalizedlichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.stack-v2-python-normal-onlyOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedmodelnet40_normal_resampled-compressedcv24-ca-128-normalizedmerged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 800,
"total_frames": 133851,
"total_tasks": 40,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:800"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task.gspc-normalized
GSPC normalised — every bank in one schema
The one schema to read first. 518 rows that flatten several GSPC banks into a
single shape: source (the bank repository the row came from), axis, category, anchor, prompt,
expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently
dropped. If you want to reuse the banks without learning each one's native layout, start here.
The live board is the authority
GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.Skill-normalisation-ESCO-graded
skill-normalisation-esco-graded
Graded-relevance annotations for surface skill terms (ESCO alt-labels) from
ESCO v1.1.0 skill-normalisation pairs
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
50
_id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.igbo_tts_normalizedv1v1d_docmatix_1k_normal_v6v1v1d_docmatix_1k_normal_v5Multilingual-Normalizer
Multilingual TTS text normalizer (written → spoken)
Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence
the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the
exact spoken form, in the same language, with nothing left that a TTS model cannot say.
52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is
digit-free on the spoken side.
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.
