datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-corpus-cleanedsmollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.ko-corpus-cleaned-12653878vi-corpus-cleaned-54988654ja-corpus-cleaned-21818123processed_book_corpus_cleanedtool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified
Nemotron-Terminal-Corpus — Cleaned & Rectified
Cleaned and restructured version of nvidia/Nemotron-Terminal-Corpus. The original dataset contains ~366K terminal agent trajectories built by NVIDIA using the Terminal-Task-Gen pipeline across math, code, SWE, and synthetic skill-based domains. This version converts the JSON-action format into a strict multi-turn conversation structure with explicit reasoning traces, validated JSON tool calls, and proper role transitions.
Original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified.macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.macedonian-corpus-cleaned
Macedonian Corpus - Cleaned
raw version here
Paper
🌟 Key Highlights
Size: 35.5 GB, Word Count: 3.31 billion
Filtered for irrelevant and low-quality content using C4 and Gopher filtering.
Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.harmonicmlx-cleaned-corpus
HarmonicMLX Cleaned Corpus v3
High-quality, balanced English text corpus for small language model pre-training.
Properly rebalanced to avoid TinyStories domination.
Pipeline
Source ingestion: FineWeb-Edu (623 MB), TinyStories (1.8 GB), Stanford Encyclopedia of Philosophy (127 MB), Project Gutenberg
Cleaning: Unicode normalization, Gutenberg/archive header stripping, URL removal, whitespace collapse
Chunking: Sentence-aware chunking (128-2048 chars)
Exact deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/harmonicmlx-cleaned-corpus.top-n-reddit-corpus-55-cleaned
Dataset Card for "top-n-reddit-corpus-55-cleaned"
More Information needed
corpus_cleanedgeo_large_corpus_cleaned_v2random-walk-reddit-corpus-55-cleaned
Dataset Card for "random-walk-reddit-corpus-55-cleaned"
More Information needed
ethos-turkish-wikipedia-cleaned-corpus-v1Cleaned_Switchboard_Corpuseecs-cleaned-corpus
EECS Cleaned Corpus
Cleaned UC Berkeley EECS crawl data exported from the CS288 assignment repository.
Contents
data/pages.jsonl: 4848 cleaned pages
data/chunks.jsonl: 22266 retrieval chunks
data/qa_live_benchmark.jsonl: 220 benchmark examples
Schema
pages.jsonl: url, title, text, page_id
chunks.jsonl: chunk-level retrieval records from the offline build pipeline
This export is produced offline from the repository's cleaned corpus artifacts.
moscar-corpus-thai-cleaned
Dataset mOSCAR Thai Cleaned
ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ
รายละเอียดชุดข้อมูล
จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train)
ขนาดข้อมูล: 5,132,779,656 ไบต์
ฟีเจอร์:
title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง
text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว
ภาษา: ไทย (th)
ลิขสิทธิ์: Apache-2.0
ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.myanmar-corpus-cleaned
Dataset Card for myanmar-corpus-cleaned
Dataset Summary
ဒီ dataset က myanmar-corpus-cleaned အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/train-00000-of-00001.parquet
validation… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-corpus-cleaned.cleaned_umls_corpussindhi-pretraining-corpus-part1-cleanedappropriateness-corpus-extension-cleanedtwi-sentiments-corpus-v1-cleanedww1-corpus-cleanedbook_corpus_cleanedKashmiri_Text_Corpus_Cleaned_2025_HNM
Overview
A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training.
Corpus Text Analysis Report
The corpus contains approximately 2 million words, with over 91,000 unique words
The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training
The cleaning process removed about 4.4% of characters while preserving 99.3% of words
Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.corpus_cleaned_testappropriateness-corpus-cleanedmyanmar_spoken_corpus_v4_cleanedCredits and Acknowledgments
This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn).
Original Dataset Source: freococo/myanmar_spoken_corpus
Modifications: * Curated and filtered for specific training needs of the ShweYon model.
Re-formatted into 36 shards for optimized Continued Pre-training (CPT).
We are deeply grateful to freococo for their contribution to the Myanmar AI community by providing this high-quality spoken corpus.
small_corpus_cleaned
