CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face02AmanPriyanshu /tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified Nemotron-Terminal-Corpus — Cleaned & Rectified Cleaned and restructured version of nvidia/Nemotron-Terminal-Corpus. The original dataset contains ~366K terminal agent trajectories built by NVIDIA using the Terminal-Task-Gen pipeline across math, code, SWE, and synthetic skill-based domains. This version converts the JSON-action format into a strict multi-turn conversation structure with explicit reasoning traces, validated JSON tool calls, and proper role transitions. Original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified.texttext-generation100K<n<1M5 likes77 downloads7mo agoHugging Face03LVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes56 downloads1y agoHugging Face04LVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes45 downloads1y agoHugging Face05Ekrem-the-second /ethos-turkish-wikipedia-cleaned-corpus-v1textquestion-answering10K<n<100K1 likes24 downloads9mo agoHugging Face06ZombitX64 /moscar-corpus-thai-cleaned Dataset mOSCAR Thai Cleaned ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ รายละเอียดชุดข้อมูล จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train) ขนาดข้อมูล: 5,132,779,656 ไบต์ ฟีเจอร์: title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว ภาษา: ไทย (th) ลิขสิทธิ์: Apache-2.0 ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.texttext-generation1M<n<10M0 likes19 downloads1y agoHugging Face07Omarrran /Kashmiri_Text_Corpus_Cleaned_2025_HNMgated Overview A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training. Corpus Text Analysis Report The corpus contains approximately 2 million words, with over 91,000 unique words The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training The cleaning process removed about 4.4% of characters while preserving 99.3% of words Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.texttext-generationn<1K0 likes7 downloads9mo agoHugging Face08rausch /scientific_corpus_cleanedgated Dataset Card for Scientific Corpus (Cleaned) This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata. Dataset Details Uses Direct Use Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.tabulartext-generation1K<n<10K2 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.