normalize
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.fishnet-lichess-normalizedXDof-TshirtFolding-20hours-normalizedperturbseq_normalized
CellClip public normalized perturbation single-cell release
This public repository contains the source-cleared, human, cell-level portion
of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and
260,859,318,437 payload bytes. These are processed derivatives rather than
the original raw download archives.
Source partition
Units
Cells
Reference cells
Effect cells
scPerturb (23 genepert + 229 chempert)
252
4,774,802
397,528
4,377,274
XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.DiffusionPDE-normalizedWe take the dataset from DiffusionPDE. For convenience, we provide our processed version on Hugging Face (see Appendix D & E in our FunDPS paper for details). The processing scripts are provided under utils/ here. It is worth noting that we normalized the datasets to zero mean and 0.5 standard deviation to follow EDM's practice.
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalized
