datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.marigold_normals_evalfishnet-lichess-normalizedXDof-TshirtFolding-20hours-normalizedT2I-ImageNet-Normalmc4-ja-filter-ja-normal
Dataset Card for "mc4-ja-filter-ja-normal"
More Information needed
router-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.oscar2301-ja-filter-ja-normal
Dataset Card for "oscar2301-ja-filter-ja-normal"
More Information needed
perturbseq_normalized
CellClip public normalized perturbation single-cell release
This public repository contains the source-cleared, human, cell-level portion
of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and
260,859,318,437 payload bytes. These are processed derivatives rather than
the original raw download archives.
Source partition
Units
Cells
Reference cells
Effect cells
scPerturb (23 genepert + 229 chempert)
252
4,774,802
397,528
4,377,274
XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.Depth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
Normalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
coco2017_caption_normalMultimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans
Multimodal Chest X ray dataset for Normal and Bacterial Pneumonia in Africans | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: imagefolder - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Multimodal-Chest-X-ray-dataset-for-Normal-and-Bacterial-Pneumonia-in-Africans.Depth-Normal-Images-617KDiffusionPDE-normalizedWe take the dataset from DiffusionPDE. For convenience, we provide our processed version on Hugging Face (see Appendix D & E in our FunDPS paper for details). The processing scripts are provided under utils/ here. It is worth noting that we normalized the datasets to zero mean and 0.5 standard deviation to follow EDM's practice.
cv24-tr-128-normalizedcv24-cy-128-normalizedOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtextures-color-normal-1k
textures-color-normal-1k
Dataset Summary
The textures-color-normal-1k dataset is an image dataset of 1000+ color and normal map textures in 512x512 resolution.
The dataset was created for use in image to image tasks.
It contains a combination of CC0 procedural and photoscanned PBR materials from ambientCG.
Dataset Structure
Data Instances
Each data point contains a 512x512 color texture and the corresponding 512x512 normal map.
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/dream-textures/textures-color-normal-1k.cv24-ur-128-normalizedcv24-pt-128-normalizedcv24-uk-128-normalizedcv24-sw-128-normalizedcv24-de-128-normalizedcv24-sk-128-normalizedlichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.stack-v2-python-normal-onlynormal_test_benchOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedmodelnet40_normal_resampled-compressed
