CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hula0401 /cad-corpus-cleanedtabular1M<n<10M4 likes13k downloads3mo agoHugging Face02Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face03anhdungitvn /ko-corpus-cleaned-12653878text10M<n<100M0 likes401 downloads3y agoHugging Face04anhdungitvn /vi-corpus-cleaned-54988654text100M<n<1B0 likes330 downloads3y agoHugging Face05anhdungitvn /ja-corpus-cleaned-21818123text10M<n<100M0 likes140 downloads3y agoHugging Face06gokulsrinivasagan /processed_book_corpus_cleaned1M<n<10M0 likes107 downloads2y agoHugging Face07AmanPriyanshu /tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified Nemotron-Terminal-Corpus — Cleaned & Rectified Cleaned and restructured version of nvidia/Nemotron-Terminal-Corpus. The original dataset contains ~366K terminal agent trajectories built by NVIDIA using the Terminal-Task-Gen pipeline across math, code, SWE, and synthetic skill-based domains. This version converts the JSON-action format into a strict multi-turn conversation structure with explicit reasoning traces, validated JSON tool calls, and proper role transitions. Original… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-CODING-Nemotron-Terminal-Corpus-data-cleaned-rectified.texttext-generation100K<n<1M5 likes82 downloads7mo agoHugging Face08LVSTCK /macedonian-corpus-cleaned-dedup Macedonian Corpus - Cleaned and Deduplicated Paper 🌟 Key Highlights Size: 16.78 GB, Word Count: 1.47 billion Deduplicated using MinHash to remove redundant documents. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated resource encompassing all available public data exists. Another challenge is the state of… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned-dedup.texttext-generation1M<n<10M1 likes57 downloads1y agoHugging Face09LVSTCK /macedonian-corpus-cleaned Macedonian Corpus - Cleaned raw version here Paper 🌟 Key Highlights Size: 35.5 GB, Word Count: 3.31 billion Filtered for irrelevant and low-quality content using C4 and Gopher filtering. Includes text from 10+ sources such as fineweb-2, HPLT-2, Wikipedia, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far as we know, no consolidated… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-cleaned.texttext-generation1M<n<10M0 likes46 downloads1y agoHugging Face10MonumentalSystems /harmonicmlx-cleaned-corpus HarmonicMLX Cleaned Corpus v3 High-quality, balanced English text corpus for small language model pre-training. Properly rebalanced to avoid TinyStories domination. Pipeline Source ingestion: FineWeb-Edu (623 MB), TinyStories (1.8 GB), Stanford Encyclopedia of Philosophy (127 MB), Project Gutenberg Cleaning: Unicode normalization, Gutenberg/archive header stripping, URL removal, whitespace collapse Chunking: Sentence-aware chunking (128-2048 chars) Exact deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/harmonicmlx-cleaned-corpus.tabular100K<n<1M0 likes45 downloads5mo agoHugging Face11dmayhem93 /top-n-reddit-corpus-55-cleaned Dataset Card for "top-n-reddit-corpus-55-cleaned" More Information needed text1M<n<10M0 likes33 downloads4y agoHugging Face12Atipico1 /corpus_cleaned0 likes31 downloads2y agoHugging Face13ZurabDz /geo_large_corpus_cleaned_v2text1M<n<10M0 likes26 downloads3y agoHugging Face14dmayhem93 /random-walk-reddit-corpus-55-cleaned Dataset Card for "random-walk-reddit-corpus-55-cleaned" More Information needed text1M<n<10M0 likes25 downloads4y agoHugging Face15Ekrem-the-second /ethos-turkish-wikipedia-cleaned-corpus-v1textquestion-answering10K<n<100K1 likes24 downloads9mo agoHugging Face16Aman279 /Cleaned_Switchboard_Corpustext1K<n<10K1 likes22 downloads3y agoHugging Face17darvog /eecs-cleaned-corpus EECS Cleaned Corpus Cleaned UC Berkeley EECS crawl data exported from the CS288 assignment repository. Contents data/pages.jsonl: 4848 cleaned pages data/chunks.jsonl: 22266 retrieval chunks data/qa_live_benchmark.jsonl: 220 benchmark examples Schema pages.jsonl: url, title, text, page_id chunks.jsonl: chunk-level retrieval records from the offline build pipeline This export is produced offline from the repository's cleaned corpus artifacts. 0 likes20 downloads6mo agoHugging Face18ZombitX64 /moscar-corpus-thai-cleaned Dataset mOSCAR Thai Cleaned ชุดข้อมูลนี้เป็นชุดข้อมูลภาษาไทยขนาดใหญ่ที่ผ่านการทำความสะอาดแล้ว เหมาะสำหรับงานประมวลผลภาษาธรรมชาติ (NLP) เช่น การฝึกสอนโมเดลภาษา การสรุปผล การแปลภาษา ฯลฯ รายละเอียดชุดข้อมูล จำนวนตัวอย่าง: 1,643,471 ตัวอย่าง (train) ขนาดข้อมูล: 5,132,779,656 ไบต์ ฟีเจอร์: title (string): หัวข้อหรือข้อความแรกของแต่ละตัวอย่าง text (string): เนื้อหาข้อความภาษาไทยที่ผ่านการคัดกรองและทำความสะอาดแล้ว ภาษา: ไทย (th) ลิขสิทธิ์: Apache-2.0 ขนาด: 1M < n < 10M… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/moscar-corpus-thai-cleaned.texttext-generation1M<n<10M0 likes18 downloads1y agoHugging Face19kkomyoeminaung /myanmar-corpus-cleaned Dataset Card for myanmar-corpus-cleaned Dataset Summary ဒီ dataset က myanmar-corpus-cleaned အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train data/train-00000-of-00001.parquet validation… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar-corpus-cleaned.text100K<n<1M0 likes18 downloads2mo agoHugging Face20Jaafer /cleaned_umls_corpustext1M<n<10M1 likes13 downloads2y agoHugging Face21jamalimubashirali /sindhi-pretraining-corpus-part1-cleanedtext1M<n<10M0 likes11 downloads6mo agoHugging Face22timonziegenbein /appropriateness-corpus-extension-cleanedtabular10K<n<100K0 likes10 downloads1y agoHugging Face23ik /twi-sentiments-corpus-v1-cleanedtext100K<n<1M0 likes9 downloads1y agoHugging Face24Fiflak1337 /ww1-corpus-cleanedtext1K<n<10K0 likes8 downloads1y agoHugging Face25BhandariShivam /book_corpus_cleanedtext10K<n<100K0 likes8 downloads8mo agoHugging Face26Omarrran /Kashmiri_Text_Corpus_Cleaned_2025_HNMgated Overview A comprehensive linear text corpus of the Kashmiri language, optimized for large language model (LLM) pre-training. Corpus Text Analysis Report The corpus contains approximately 2 million words, with over 91,000 unique words The entire corpus is formatted as a single continuous line of text, ideal for LLM Pre-training The cleaning process removed about 4.4% of characters while preserving 99.3% of words Vocabulary richness… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Kashmiri_Text_Corpus_Cleaned_2025_HNM.texttext-generationn<1K0 likes7 downloads9mo agoHugging Face27Atipico1 /corpus_cleaned_test0 likes4 downloads2y agoHugging Face28timonziegenbein /appropriateness-corpus-cleanedtabularn<1K0 likes4 downloads1y agoHugging Face29URajinda /myanmar_spoken_corpus_v4_cleanedCredits and Acknowledgments This dataset is built upon the foundational work of the Myanmar Spoken Corpus by freococo (Wynn). Original Dataset Source: freococo/myanmar_spoken_corpus Modifications: * Curated and filtered for specific training needs of the ShweYon model. Re-formatted into 36 shards for optimized Continued Pre-training (CPT). We are deeply grateful to freococo for their contribution to the Myanmar AI community by providing this high-quality spoken corpus. text1M<n<10M0 likes4 downloads8mo agoHugging Face30ashercn97 /small_corpus_cleanedtabular10K<n<100K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.