datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-culturax-clean-dataset
Thai CulturaX Clean dataset
The data is sourced from the Thai subset of CulturaX dataset, which itself is sourced from mC4 and four OSCAR corpora.
It has about 8,748,575,684 words (without whitespace) and 16,768,585 lines (97 GB).
It was filtered content promoting gambling, adult content, and narcotics.
GitHub for clean: https://github.com/wannaphong/thai-filter-website
Considerations for Using the Data
This dataset is the cleaned version of the CulturaX datasets… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-culturax-clean-dataset.gutenberg-en-v1-clean
gutenberg - clean
dataset_info:
- config_name: default
features:
- name: text
dtype: string
- name: label
dtype: string
- name: score
dtype: float64
- name: sha256dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 3384868097
num_examples: 9978
- name: validation
num_bytes: 195405579
num_examples: 574
- name: test
num_bytes: 189439446
num_examples: 565
download_size: 2317462261
dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.clean-songs-lyrics-dataset
Clean Songs Lyrics Dataset
1.53M+ clean songs lyrics with songs titles and artists names
Dataset info
This is the combined, deduped, cleaned, and sanitized aggregation of three large lyrics datasets
Each lyric was deduplicated
Each lyric was checked to be in range of 256 bytes <-> 8192 bytes
Each lyric was checked for profanities with alt-profanity-check
Each lyric was ASCII sanitized for conistency… See the full description on the dataset page: https://huggingface.co/datasets/asigalov61/clean-songs-lyrics-dataset.clean-tr-dataset
Turkish Clean Text Corpus
Dataset Description
This dataset is a highly cleaned, deduplicated, and normalized corpus of Turkish text. It features a diverse collection of informative, encyclopedic, and journalistic content.
The dataset is primarily designed for Continued Pre-Training (CPT) of Large Language Models (LLMs) to enhance their Turkish language capabilities and domain knowledge. It is also an excellent foundational corpus for generating synthetic… See the full description on the dataset page: https://huggingface.co/datasets/mehmettozlu/clean-tr-dataset.merged_dataset_final_clean_v41
merged_dataset_final_clean_v41
English
Rule-based cleaned SFT dataset for structured output generation (JSON / YAML / XML / TOML / CSV).
Data Source
This dataset was built from competition-provided datasets only.
The cleaning pipeline loads the following source groups:
u-10bei (6 datasets: source ids 1-1 to 1-6)
daichira (3 datasets: source ids 2-1 to 2-3)
After strict filtering and sampling for v4.1, the final retained rows are from u-10bei sources (1-1 to… See the full description on the dataset page: https://huggingface.co/datasets/yamaTK/merged_dataset_final_clean_v41.Clean-Alignment-Dataset
Clean Alignment Dataset
What is this dataset?
Clean Alignment Dataset is a safety preference dataset for Direct Preference
Optimization (DPO) and related preference-alignment methods. Every example is a
(prompt, chosen, rejected) triple in which the chosen response is safe and
the rejected response is unsafe for the same prompt — an unambiguous,
consistently-labelled safe-vs-unsafe contrast in every single pair.
It is built by combining and re-cleaning two… See the full description on the dataset page: https://huggingface.co/datasets/etrigan5500/Clean-Alignment-Dataset.
