CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B45 likes176k downloads3mo agoHugging Face02argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face03Hula0401 /cad-corpus-cleanedtabular1M<n<10M4 likes13k downloads3mo agoHugging Face04Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes2.9k downloads9mo agoHugging Face05kejian /codeparrot-train-more-filter-3.3b-cleanedtabulartext-classification1M<n<10M2 likes2.3k downloads4y agoHugging Face06allenai /ultrafeedback_binarized_cleaned Dataset Card for "ultrafeedback_binarized_cleaned" Update 1/12/2023: I've removed examples identified as faulty by Argilla - see their awesome work for more details. This is a version of the UltraFeedback binarized dataset but with TruthfulQA prompts removed and source annotations added (so you can filter out samples from different sources yourself if you want!). Please see the binarized dataset card for more information, or the original UltraFeedback dataset card. tabular100K<n<1M72 likes1.7k downloads3y agoHugging Face07moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads26d agoHugging Face08dacorvo /funes-xiaowu0162-longmemeval-cleaned-s Funes recall store — LongMemEval_s cleaned corpus A funes recall store built by indexing the longmemeval_s_cleaned.json haystack of xiaowu0162/longmemeval-cleaned (LongMemEval, arXiv:2410.10813) — every unique chat session across all 500 questions' haystacks, in one corpus-wide store. What this is This is not a raw trace dataset — it is a pre-built funes index: the source sessions chunked into content blocks and embedded, stored as a Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.tabular100K<n<1M0 likes1.4k downloads2mo agoHugging Face09Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.4k downloads3y agoHugging Face10tiagoloeblein /CrawlPT_dedup_Cleaned📚 CrawlPT Clean — High-Quality Portuguese Corpus Versão limpa, filtrada e refinada do dataset CrawlPT_dedup 🧼 Visão Geral Este repositório fornece uma versão limpa, filtrada e padronizada do dataset: ➡️ eduagarcia/CrawlPT_dedup https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup A limpeza tem como objetivo criar um corpus de alta qualidade para: pré-treino contínuo de modelos LLM (Qwen, Mistral, LLaMA, Phi etc.) melhora de fluência e coerência em português pesquisas em NLP geração de… See the full description on the dataset page: https://huggingface.co/datasets/tiagoloeblein/CrawlPT_dedup_Cleaned.tabular100M<n<1B1 likes783 downloads10mo agoHugging Face11Finnish-NLP /oscar_2301_fi_cleaned Dataset Card for "oscar_2301_fi_cleaned" More Information needed tabular1M<n<10M0 likes763 downloads3y agoHugging Face12Avelina /python-edu-cleaned SmolLM-Corpus: Python-Edu (Cleaned) This dataset contains the python-edu subset of SmolLM-Corpus with the contents of the files stored in a new text field. All files were downloaded from the S3 bucket on January the 8th 2025, using the blob IDs from the original dataset with revision 3ba9d605774198c5868892d7a8deda78031a781f. Only 1 file was marked as not found and the corresponding row removed from the dataset (content/39c3e5b85cc678d1d54b4d93a55271c51d54126c which I suspect is… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/python-edu-cleaned.tabular1M<n<10M3 likes756 downloads2y agoHugging Face13oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes608 downloads23d agoHugging Face14VibeCuisine /jetson1-060926-subtask-place-full-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place-full-cleaned.tabularrobotics1K<n<10K0 likes600 downloads3mo agoHugging Face15VibeCuisine /jetson1-060826-subtask-grab2-full-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2-full-cleaned.tabularrobotics1K<n<10K0 likes591 downloads3mo agoHugging Face16dd-n-kk /uci-drug-review-cleanedtabular100K<n<1M0 likes577 downloads2y agoHugging Face17ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes550 downloads3mo agoHugging Face18KeisukeMiyamoto /CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset. This dataset was created from the sample_10BT subset of hotchpotch/fineweb-2-edu-japanese. The source text was refined with MK0727/corpus-refiner-jp. Purpose The main purpose of this dataset is to provide cleaner Japanese web text for language model pretraining and continued pretraining. This dataset keeps Japanese web documents from FineWeb2-Edu while reducing boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.tabulartext-generation10M<n<100M1 likes362 downloads1mo agoHugging Face19pszemraj /qmsum-cleaned qmsum-cleaned prefixes It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset. Most frequent "prefixes" separated via sentence-splitter in the train split: Sentence Count 0 Summarize the whole meeting. 121 1 Summarize the meeting 25 2 What did the team discuss about the product cost? 4 3 How did… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.tabularsummarization1K<n<10K14 likes324 downloads9mo agoHugging Face20DanielTobi0 /openresearcher-sft-deep-research-cleaned OpenResearcher SFT DeepResearch — Parquet Mirror This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab. Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.tabulartext-generation10K<n<100K0 likes283 downloads2mo agoHugging Face21formalmathatepfl /sft-one_shot-cleanedtabular1M<n<10M0 likes274 downloads25d agoHugging Face22BramVanroy /ultra_feedback_dutch_cleaned Ultra Feedback Dutch Cleaned This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used. After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.tabulartext-generation100K<n<1M6 likes259 downloads2y agoHugging Face23KeisukeMiyamoto /CleanedWiki-jp CleanedWiki-jp CleanedWiki-jp is a cleaned Japanese Wikipedia dataset prepared for LLM pre-training. It is built from Japanese Wikipedia article HTML, converted into Markdown, filtered for trainability. The dataset keeps useful article structure instead of flattening everything into plain text. Suitable body tables are preserved as Markdown tables, and mathematical expressions are preserved in TeX form. Each row also includes a predicted Nippon Decimal Classification (NDC)… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedWiki-jp.tabulartext-generation1M<n<10M0 likes241 downloads1mo agoHugging Face24imoxto /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M6 likes228 downloads3y agoHugging Face25it4lia /EMBER_cleaned EMBER Cleaned EMBER Cleaned is a cleaned and AI-ready version of the original EMBER (Endgame Malware Benchmark for Research) dataset, a widely used benchmark for static malware detection on Windows Portable Executable (PE) files. The original EMBER dataset was introduced by Endgame / Elastic as an open benchmark for machine-learning-based malware detection using only static PE-derived features, without executing binaries. This cleaned release preserves that purpose while making the… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/EMBER_cleaned.tabulartabular-classification100K<n<1M0 likes201 downloads6mo agoHugging Face26zerostratos /fineweb-2-vie-2022-cleanedtabular1M<n<10M0 likes198 downloads1y agoHugging Face27myzxyz /my_dataset_cleaned my_dataset This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. 数据集信息 总episodes数: 47 (原48个,已移除episode_000000) 总帧数: 8,918 任务: 抓取立方体并放入盒子 机器人: so-100 帧率: 30 FPS 数据质量说明 注意: 原始数据集中的第0个episode (episode_000000) 由于视频质量问题已被移除。当前数据集从episode_000001开始,包含47个高质量的episode。… See the full description on the dataset page: https://huggingface.co/datasets/myzxyz/my_dataset_cleaned.tabularrobotics1K<n<10K0 likes190 downloads1y agoHugging Face28anothy1 /fineweb-edu-cleaned-simplifiedtabular10K<n<100K2 likes185 downloads2y agoHugging Face29omid5 /usda-fdc-foods-cleaned Comprehensive & Cleaned USDA Foods Nutrition Dataset Dataset Summary This dataset is a cleaned, de-duplicated, and enhanced version of the USDA's FoodData Central (FDC) database, combining Branded Foods, Foundation Foods (generic), and SR Legacy data into a single, analysis-ready file. It is designed to be a robust resource for nutritional analysis, machine learning, and food-related applications. The raw USDA data is spread across dozens of CSV files, contains numerous… See the full description on the dataset page: https://huggingface.co/datasets/omid5/usda-fdc-foods-cleaned.tabular100K<n<1M2 likes181 downloads1y agoHugging Face30jayp132 /green-only-200-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/jayp132/green-only-200-cleaned.tabularrobotics10K<n<100K0 likes179 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.