datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedcv24-tr-128-normalizedcv24-cy-128-normalizedcv24-ur-128-normalizedcv24-uk-128-normalizedcv24-de-128-normalizedcv24-pt-128-normalizedcv24-sw-128-normalizedNormalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedcv24-sk-128-normalizedlichess-stockfish-normalized
Lichess Chess Positions: ML-Ready Deduplicated Evaluations
Dataset Description
A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database.
Why This Dataset?
While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers:
Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.cv24-ca-128-normalizedgspc-normalized
GSPC normalised — every bank in one schema
The one schema to read first. 518 rows that flatten several GSPC banks into a
single shape: source (the bank repository the row came from), axis, category, anchor, prompt,
expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently
dropped. If you want to reuse the banks without learning each one's native layout, start here.
The live board is the authority
GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.igbo_tts_normalizedVFD_normalize_9_v1meld-open-normalized
MELD Open (Normalized)
MELD is a multilingual and multi-domain dataset for Named Entity Recognition (NER) constructed from 60 existing datasets. It includes gold-standard annotations across 60 languages and 14 domains. This dataset is a subset of 43 datasets for which licenses permit the redistribution of data in a new format. See the MELD GitHub repository for more details.
Note: This version of MELD Open uses normalized labels. For original labels from each source dataset, use… See the full description on the dataset page: https://huggingface.co/datasets/kgnlp/meld-open-normalized.minif2f-lean4-normalizedtask093_conala_normalize_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task093_conala_normalize_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task093_conala_normalize_lists.cv24-ru-128-normalizedFleurs_Irish_normalizedcorpus-dataset-normalized-for-persian-and-english
Dataset Summary
Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks.
To see Persian data in Viewer tab click here
English data of this dataset is merged from english-wiki-corpus dataset.
Note: If you need only Persian corpus click here
Note: The data for both Persian… See the full description on the dataset page: https://huggingface.co/datasets/ali619/corpus-dataset-normalized-for-persian-and-english.Arabic-Diacritized-TTS-Normalized
Arabic-Diacritized-TTS Dataset
Overview
The Arabic-Diacritized-TTS dataset contains Arabic audio samples and their corresponding text with full diacritization. This dataset is designed to support research in Arabic speech processing, text-to-speech (TTS) synthesis, automatic diacritization, and other natural language processing (NLP) tasks.
Dataset Contents
Audio Samples: High-quality Arabic speech recordings.
Text Transcriptions: Fully diacritized Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/hana92/Arabic-Diacritized-TTS-Normalized.common-voice-20-mn-normalized
Common Voice 20.0 Mongolian Dataset
This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release.
Dataset Structure
The dataset contains:
Audio clips in .mp3 format
Transcriptions for each audio clip
Train/test/dev splits
Additional metadata including speaker demographics
Usage
This dataset can be used for:
Speech Recognition
Voice Analysis
Linguistic Research
Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.corpus-dataset-normalized-for-persian-farsi
Dataset Summary
Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks.
The data in this dataset have been normalized and unnecessary tokens have been removed.
Note: If you need Persian and Engish corpus together, click here
OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk256-normalizedmasc_filtered_normalizedMalaysian-Normalizer
Malaysian Normalizer
Normalize numbers, digits, currency, IC, time, date, timestamp, email, URL, titles, abbrevations and symbols.
Instruction format
We make sure the normalized text able to reverse back to the original text using the normalized mapping.
We make sure the normalized text does not contain any digits.
We predict major language in the normalized mapping and use it as prompt language.
Convert to instruction format, uploaded at… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Normalizer.mega-reasoning-mix-normalized-v3
🧠 Mega Reasoning Mix Normalized v3
Dataset Summary
The Mega Reasoning Mix Normalized v3 is a highly curated, large-scale dataset designed for fine-tuning Large Language Models (LLMs) on complex reasoning, step-by-step Chain-of-Thought (CoT), advanced coding tasks, and agentic tool use.
This dataset is the result of combining over a dozen top-tier synthetic and filtered reasoning datasets. The entire corpus has been strictly normalized into a unified schema and… See the full description on the dataset page: https://huggingface.co/datasets/dschauhan08/mega-reasoning-mix-normalized-v3.VFD_normalize_9
