CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.7k downloads4mo agoHugging Face02SaakethS /fishnet-lichess-normalized1 likes2.1k downloads11d agoHugging Face03TongheZhangTH /XDof-TshirtFolding-20hours-normalizedtabular1M<n<10M0 likes1.5k downloads8mo agoHugging Face04cfy2yue /perturbseq_normalized CellClip public normalized perturbation single-cell release This public repository contains the source-cleared, human, cell-level portion of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and 260,859,318,437 payload bytes. These are processed derivatives rather than the original raw download archives. Source partition Units Cells Reference cells Effect cells scPerturb (23 genepert + 229 chempert) 252 4,774,802 397,528 4,377,274 XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.feature-extraction0 likes863 downloads2mo agoHugging Face05jcy20 /DiffusionPDE-normalizedWe take the dataset from DiffusionPDE. For convenience, we provide our processed version on Hugging Face (see Appendix D & E in our FunDPS paper for details). The processing scripts are provided under utils/ here. It is worth noting that we normalized the datasets to zero mean and 0.5 standard deviation to follow EDM's practice. 10K<n<100K2 likes545 downloads1y agoHugging Face06winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes529 downloads1y agoHugging Face07N03N9 /cv24-tr-128-normalizedtext100K<n<1M0 likes526 downloads9mo agoHugging Face08N03N9 /cv24-cy-128-normalizedtext10K<n<100K0 likes487 downloads9mo agoHugging Face09N03N9 /cv24-ur-128-normalizedtext10K<n<100K0 likes409 downloads9mo agoHugging Face10N03N9 /cv24-uk-128-normalizedtext10K<n<100K0 likes400 downloads9mo agoHugging Face11N03N9 /cv24-de-128-normalizedtext100K<n<1M0 likes397 downloads9mo agoHugging Face12N03N9 /cv24-pt-128-normalizedtext100K<n<1M0 likes389 downloads9mo agoHugging Face13N03N9 /cv24-sw-128-normalizedtext100K<n<1M0 likes385 downloads9mo agoHugging Face14Scicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes382 downloads6mo agoHugging Face15winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes374 downloads1y agoHugging Face16N03N9 /cv24-sk-128-normalizedtext10K<n<100K0 likes355 downloads9mo agoHugging Face17mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes343 downloads10mo agoHugging Face18N03N9 /cv24-ca-128-normalizedtext1M<n<10M0 likes306 downloads9mo agoHugging Face19jxie /epsilon-normalized Dataset Card for "epsilon-normalized" More Information needed 100K<n<1M0 likes282 downloads3y agoHugging Face20csoai /gspc-normalized GSPC normalised — every bank in one schema The one schema to read first. 518 rows that flatten several GSPC banks into a single shape: source (the bank repository the row came from), axis, category, anchor, prompt, expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently dropped. If you want to reuse the banks without learning each one's native layout, start here. The live board is the authority GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.textothern<1K0 likes254 downloads11d agoHugging Face21Twelve2five /igbo_tts_normalizedaudio100K<n<1M2 likes240 downloads1y agoHugging Face22Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes207 downloads20d agoHugging Face23Kudod /VFD_normalize_9_v1tabular100K<n<1M0 likes167 downloads3mo agoHugging Face24kgnlp /meld-open-normalized MELD Open (Normalized) MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details. Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.tabulartoken-classification10M<n<100M0 likes165 downloads5mo agoHugging Face25pkuAI4M /minif2f-lean4-normalizedtextn<1K4 likes161 downloads2y agoHugging Face26Lots-of-LoRAs /task093_conala_normalize_lists Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task093_conala_normalize_lists Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task093_conala_normalize_lists.texttext-generation1K<n<10K0 likes157 downloads2y agoHugging Face27jxie /qg-tagging-normalized Dataset Card for "qg-tagging-normalized" More Information needed 1M<n<10M0 likes153 downloads3y agoHugging Face28N03N9 /cv24-ru-128-normalizedtext100K<n<1M0 likes151 downloads9mo agoHugging Face29Eimhin03 /Fleurs_Irish_normalizedaudio1K<n<10K0 likes144 downloads6mo agoHugging Face30ali619 /corpus-dataset-normalized-for-persian-and-english Dataset Summary Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks. To see Persian data in Viewer tab click here English data of this dataset is merged from english-wiki-corpus dataset. Note: If you need only Persian corpus click here Note: The data for both Persian… See the full description on the dataset page: https://huggingface.co/datasets/ali619/corpus-dataset-normalized-for-persian-and-english.text1M<n<10M1 likes142 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.