datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified
Nemotron-Terminal-Corpus — Cleaned & Rectified
Cleaned and restructured version of nvidia/Nemotron-Terminal-Corpus. The original dataset contains ~366K terminal agent trajectories built by NVIDIA using the Terminal-Task-Gen pipeline across math, code, SWE, and synthetic skill-based domains. This version converts the JSON-action format into a strict multi-turn conversation structure with explicit reasoning traces, validated JSON tool calls, and proper role transitions.
Original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified.macedonian-corpus-cleaned-dedup
Macedonian Corpus - Cleaned and Deduplicated
Paper
🌟 Key Highlights
Size: 16.78 GB, Word Count: 1.47 billion
Deduplicated using MinHash to remove redundant documents.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.macedonian-corpus-cleaned
Macedonian Corpus - Cleaned
raw version here
Paper
🌟 Key Highlights
Size: 35.5 GB, Word Count: 3.31 billion
Filtered for irrelevant and low-quality content using C4 and Gopher filtering.
Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more.
📋 Overview
Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.ethos-turkish-wikipedia-cleaned-corpus-v1moscar-corpus-thai-cleaned
Dataset mOSCAR Thai Cleaned
ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ
รายละเอียดชุดข้อมูล
จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train)
ขนาดข้อมูล: 5,132,779,656 ไบต์
ฟีเจอร์:
title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง
text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว
ภาษา: ไทย (th)
ลิขสิทธิ์: Apache-2.0
ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.Kashmiri_Text_Corpus_Cleaned_2025_HNM
Overview
A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training.
Corpus Text Analysis Report
The corpus contains approximately 2 million words, with over 91,000 unique words
The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training
The cleaning process removed about 4.4% of characters while preserving 99.3% of words
Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.scientific_corpus_cleaned
Dataset Card for Scientific Corpus (Cleaned)
This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata.
Dataset Details
Uses
Direct Use
Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.
