CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes529 downloads1y agoHugging Face02N03N9 /cv24-tr-128-normalizedtext100K<n<1M0 likes526 downloads9mo agoHugging Face03N03N9 /cv24-cy-128-normalizedtext10K<n<100K0 likes487 downloads9mo agoHugging Face04N03N9 /cv24-ur-128-normalizedtext10K<n<100K0 likes409 downloads9mo agoHugging Face05N03N9 /cv24-uk-128-normalizedtext10K<n<100K0 likes400 downloads9mo agoHugging Face06N03N9 /cv24-de-128-normalizedtext100K<n<1M0 likes397 downloads9mo agoHugging Face07N03N9 /cv24-pt-128-normalizedtext100K<n<1M0 likes389 downloads9mo agoHugging Face08N03N9 /cv24-sw-128-normalizedtext100K<n<1M0 likes385 downloads9mo agoHugging Face09Scicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes382 downloads6mo agoHugging Face10winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes374 downloads1y agoHugging Face11N03N9 /cv24-sk-128-normalizedtext10K<n<100K0 likes355 downloads9mo agoHugging Face12mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes343 downloads10mo agoHugging Face13N03N9 /cv24-ca-128-normalizedtext1M<n<10M0 likes306 downloads9mo agoHugging Face14csoai /gspc-normalized GSPC normalised — every bank in one schema The one schema to read first. 518 rows that flatten several GSPC banks into a single shape: source (the bank repository the row came from), axis, category, anchor, prompt, expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently dropped. If you want to reuse the banks without learning each one's native layout, start here. The live board is the authority GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.textothern<1K0 likes254 downloads11d agoHugging Face15Twelve2five /igbo_tts_normalizedaudio100K<n<1M2 likes240 downloads1y agoHugging Face16Kudod /VFD_normalize_9_v1tabular100K<n<1M0 likes167 downloads4mo agoHugging Face17kgnlp /meld-open-normalized MELD Open (Normalized) MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details. Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.tabulartoken-classification10M<n<100M0 likes165 downloads5mo agoHugging Face18pkuAI4M /minif2f-lean4-normalizedtextn<1K4 likes161 downloads2y agoHugging Face19Lots-of-LoRAs /task093_conala_normalize_lists Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task093_conala_normalize_lists Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task093_conala_normalize_lists.texttext-generation1K<n<10K0 likes157 downloads2y agoHugging Face20N03N9 /cv24-ru-128-normalizedtext100K<n<1M0 likes151 downloads9mo agoHugging Face21Eimhin03 /Fleurs_Irish_normalizedaudio1K<n<10K0 likes144 downloads6mo agoHugging Face22ali619 /corpus-dataset-normalized-for-persian-and-english Dataset Summary Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks. To see Persian data in Viewer tab click here English data of this dataset is merged from english-wiki-corpus dataset. Note: If you need only Persian corpus click here Note: The data for both Persian… See the full description on the dataset page: https://huggingface.co/datasets/ali619/corpus-dataset-normalized-for-persian-and-english.text1M<n<10M1 likes142 downloads2y agoHugging Face23hana92 /Arabic-Diacritized-TTS-Normalized Arabic-Diacritized-TTS Dataset Overview The Arabic-Diacritized-TTS dataset contains Arabic audio samples and their corresponding text with full diacritization. This dataset is designed to support research in Arabic speech processing, text-to-speech (TTS) synthesis, automatic diacritization, and other natural language processing (NLP) tasks. Dataset Contents Audio Samples: High-quality Arabic speech recordings. Text Transcriptions: Fully diacritized Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/hana92/Arabic-Diacritized-TTS-Normalized.audio1K<n<10K0 likes119 downloads7mo agoHugging Face24warmestman /common-voice-20-mn-normalized Common Voice 20.0 Mongolian Dataset This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release. Dataset Structure The dataset contains: Audio clips in .mp3 format Transcriptions for each audio clip Train/test/dev splits Additional metadata including speaker demographics Usage This dataset can be used for: Speech Recognition Voice Analysis Linguistic Research Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.audioautomatic-speech-recognition10K<n<100K4 likes109 downloads2y agoHugging Face25ali619 /corpus-dataset-normalized-for-persian-farsi Dataset Summary Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks. The data in this dataset have been normalized and unnecessary tokens have been removed. Note: If you need Persian and Engish corpus together, click here text100K<n<1M5 likes98 downloads2y agoHugging Face26winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk256-normalizedtabular10K<n<100K0 likes98 downloads1y agoHugging Face27omarabb315 /masc_filtered_normalizedaudio100K<n<1M0 likes98 downloads1y agoHugging Face28mesolitica /Malaysian-Normalizer Malaysian Normalizer Normalize numbers, digits, currency, IC, time, date, timestamp, email, URL, titles, abbrevations and symbols. Instruction format We make sure the normalized text able to reverse back to the original text using the normalized mapping. We make sure the normalized text does not contain any digits. We predict major language in the normalized mapping and use it as prompt language. Convert to instruction format, uploaded at… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Normalizer.text1M<n<10M0 likes97 downloads1y agoHugging Face29dschauhan08 /mega-reasoning-mix-normalized-v3 🧠 Mega Reasoning Mix Normalized v3 Dataset Summary The Mega Reasoning Mix Normalized v3 is a highly curated, large-scale dataset designed for fine-tuning Large Language Models (LLMs) on complex reasoning, step-by-step Chain-of-Thought (CoT), advanced coding tasks, and agentic tool use. This dataset is the result of combining over a dozen top-tier synthetic and filtered reasoning datasets. The entire corpus has been strictly normalized into a unified schema and… See the full description on the dataset page: https://huggingface.co/datasets/dschauhan08/mega-reasoning-mix-normalized-v3.text100K<n<1M1 likes92 downloads6mo agoHugging Face30Kudod /VFD_normalize_9tabular10K<n<100K0 likes85 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.