CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face02Hula0401 /cad-corpus-cleanedtabular1M<n<10M4 likes15k downloads4mo agoHugging Face03Dr3dre /Genius-song-lyrics-cleaned 🎵 Genius Song Lyrics cleaned Dataset Dataset Description This dataset is originally taken from Genius Song Lyrics and it contains cleaned and normalized song lyrics for more than 5 million songs, designed for large-scale topic modeling, clustering, and semantic analysis. The dataset was specifically preprocessed to be compatible with embedding-based models (e.g. Sentence Transformers, BERTopic) while preserving lyrical meaning and thematic content. Repetitive structures… See the full description on the dataset page: https://huggingface.co/datasets/Dr3dre/Genius-song-lyrics-cleaned.tabulartext-classification1M<n<10M5 likes4.1k downloads9mo agoHugging Face04kejian /codeparrot-train-more-filter-3.3b-cleanedtabulartext-classification1M<n<10M2 likes2.7k downloads4y agoHugging Face05moganai /turkishfineweb2-cleaned TurkishFineweb2-Cleaned A Turkish web corpus derived from the Turkish (tur_Latn) subset of FineWeb-2, augmented with an additional quality-classification layer and a near-duplicate removal pass. 📄 Paper: MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM→MLM Curriculum Source FineWeb-2 is a large-scale, multilingual web corpus built from Common Crawl. This dataset covers the Turkish (tur_Latn) portion of FineWeb-2, spanning the… See the full description on the dataset page: https://huggingface.co/datasets/moganai/turkishfineweb2-cleaned.tabulartext-generation10M<n<100M5 likes1.7k downloads2d agoHugging Face06allenai /ultrafeedback_binarized_cleaned Dataset Card for "ultrafeedback_binarized_cleaned" Update 1/12/2023: I've removed examples identified as faulty by Argilla - see their awesome work for more details. This is a version of the UltraFeedback binarized dataset but with TruthfulQA prompts removed and source annotations added (so you can filter out samples from different sources yourself if you want!). Please see the binarized dataset card for more information, or the original UltraFeedback dataset card. tabular100K<n<1M72 likes1.7k downloads3y agoHugging Face07Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.4k downloads3y agoHugging Face08tiagoloeblein /CrawlPT_dedup_Cleaned📚 CrawlPT Clean — High-Quality Portuguese Corpus Versão limpa, filtrada e refinada do dataset CrawlPT_dedup 🧼 Visão Geral Este repositório fornece uma versão limpa, filtrada e padronizada do dataset: ➡️ eduagarcia/CrawlPT_dedup https://huggingface.co/datasets/eduagarcia/CrawlPT_dedup A limpeza tem como objetivo criar um corpus de alta qualidade para: pré-treino contínuo de modelos LLM (Qwen, Mistral, LLaMA, Phi etc.) melhora de fluência e coerência em português pesquisas em NLP geração de… See the full description on the dataset page: https://huggingface.co/datasets/tiagoloeblein/CrawlPT_dedup_Cleaned.tabular100M<n<1B1 likes951 downloads10mo agoHugging Face09Avelina /python-edu-cleaned SmolLM-Corpus: Python-Edu (Cleaned) This dataset contains the python-edu subset of SmolLM-Corpus with the contents of the files stored in a new text field. All files were downloaded from the S3 bucket on January the 8th 2025, using the blob IDs from the original dataset with revision 3ba9d605774198c5868892d7a8deda78031a781f. Only 1 file was marked as not found and the corresponding row removed from the dataset (content/39c3e5b85cc678d1d54b4d93a55271c51d54126c which I suspect is… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/python-edu-cleaned.tabular1M<n<10M3 likes790 downloads2y agoHugging Face10dd-n-kk /uci-drug-review-cleanedtabular100K<n<1M0 likes758 downloads2y agoHugging Face11Finnish-NLP /oscar_2301_fi_cleaned Dataset Card for "oscar_2301_fi_cleaned" More Information needed tabular1M<n<10M0 likes690 downloads3y agoHugging Face12oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes681 downloads26d agoHugging Face13VibeCuisine /jetson1-060926-subtask-place-full-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060926-subtask-place-full-cleaned.tabularrobotics1K<n<10K0 likes676 downloads3mo agoHugging Face14VibeCuisine /jetson1-060826-subtask-grab2-full-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-060826-subtask-grab2-full-cleaned.tabularrobotics1K<n<10K0 likes665 downloads3mo agoHugging Face15KeisukeMiyamoto /CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp CleanedFineWeb2Edu-jp is a cleaned Japanese web text dataset. This dataset was created from the sample_10BT subset of hotchpotch/fineweb-2-edu-japanese. The source text was refined with MK0727/corpus-refiner-jp. Purpose The main purpose of this dataset is to provide cleaner Japanese web text for language model pretraining and continued pretraining. This dataset keeps Japanese web documents from FineWeb2-Edu while reducing boilerplate… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedFineWeb2Edu-jp.tabulartext-generation10M<n<100M1 likes380 downloads1mo agoHugging Face16pszemraj /qmsum-cleaned qmsum-cleaned prefixes It's worth noting that each "document" in input is prefixed by a question/prompt on what the model is supposed to do. You may want to explicitly handle this in some way, or prefix your models trained on this dataset. Most frequent "prefixes" separated via sentence-splitter in the train split: Sentence Count 0 Summarize the whole meeting. 121 1 Summarize the meeting 25 2 What did the team discuss about the product cost? 4 3 How did… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/qmsum-cleaned.tabularsummarization1K<n<10K14 likes289 downloads9mo agoHugging Face17formalmathatepfl /sft-one_shot-cleanedtabular1M<n<10M0 likes279 downloads28d agoHugging Face18KeisukeMiyamoto /CleanedWiki-jp CleanedWiki-jp CleanedWiki-jp is a cleaned Japanese Wikipedia dataset prepared for LLM pre-training. It is built from Japanese Wikipedia article HTML, converted into Markdown, filtered for trainability. The dataset keeps useful article structure instead of flattening everything into plain text. Suitable body tables are preserved as Markdown tables, and mathematical expressions are preserved in TeX form. Each row also includes a predicted Nippon Decimal Classification (NDC)… See the full description on the dataset page: https://huggingface.co/datasets/KeisukeMiyamoto/CleanedWiki-jp.tabulartext-generation1M<n<10M0 likes256 downloads1mo agoHugging Face19DanielTobi0 /openresearcher-sft-deep-research-cleaned OpenResearcher SFT DeepResearch — Parquet Mirror This is a re-hosted copy of the tool-reasoning SFT deep-research dataset by Aman Priyanshu, itself a cleaned/restructured version of the OpenResearcher Dataset from TIGER-AI-Lab. Why this repo exists: the source wasn't laid out as ready-to-download Parquet files. This mirror simply stores the data as plain seed_*.parquet files so you can grab the whole dataset or a single segment easily. No changes were made to the content — all… See the full description on the dataset page: https://huggingface.co/datasets/DanielTobi0/openresearcher-sft-deep-research-cleaned.tabulartext-generation10K<n<100K0 likes254 downloads2mo agoHugging Face20zerostratos /fineweb-2-vie-2022-cleanedtabular1M<n<10M0 likes252 downloads1y agoHugging Face21BramVanroy /ultra_feedback_dutch_cleaned Ultra Feedback Dutch Cleaned This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used. After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.tabulartext-generation100K<n<1M6 likes231 downloads2y agoHugging Face22imoxto /prompt_injection_cleaned_dataset Dataset Card for "prompt_injection_cleaned_dataset" More Information needed tabular100K<n<1M6 likes218 downloads3y agoHugging Face23lelouch0204 /cleaned_allsides_v2.csvtabular1K<n<10K0 likes197 downloads11mo agoHugging Face24omid5 /usda-fdc-foods-cleaned Comprehensive & Cleaned USDA Foods Nutrition Dataset Dataset Summary This dataset is a cleaned, de-duplicated, and enhanced version of the USDA's FoodData Central (FDC) database, combining Branded Foods, Foundation Foods (generic), and SR Legacy data into a single, analysis-ready file. It is designed to be a robust resource for nutritional analysis, machine learning, and food-related applications. The raw USDA data is spread across dozens of CSV files, contains numerous… See the full description on the dataset page: https://huggingface.co/datasets/omid5/usda-fdc-foods-cleaned.tabular100K<n<1M2 likes191 downloads1y agoHugging Face25jayp132 /green-only-200-cleanedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/jayp132/green-only-200-cleaned.tabularrobotics10K<n<100K0 likes180 downloads20d agoHugging Face26anothy1 /fineweb-edu-cleaned-simplifiedtabular10K<n<100K2 likes176 downloads2y agoHugging Face27maneshkarun /hyperpartisan-cleanedtabular100K<n<1M0 likes161 downloads3y agoHugging Face28astro-legacy-archive /swift-xrt-cleaned-events Swift XRT cleaned events This dataset contains the Swift XRT cleaned photon-counting-mode event file sw00020000001xpcw4po_cl.evt.gz for public observation 00020000001, target GRB041217. The source is sw00020000001xpcw4po_cl.evt.gz, checked 2026-08-30: 47,326 bytes, SHA-256 65c8bd98bd37a184ee584459a25c09e9c0af92b0cc8235eb5071d593a2d3e61c. python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/swift-xrt-cleaned-events.tabularn<1K0 likes161 downloads20d agoHugging Face29argilla /ultrafeedback-multi-binarized-preferences-cleaned UltraFeedback - Multi-Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences-cleaned, and has been created to explore whether DPO fine-tuning with more than one rejection per chosen response helps the model perform better in the AlpacaEval, MT-Bench, and LM Eval Harness benchmarks. Read more about Argilla's approach towards UltraFeedback binarization at… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-multi-binarized-preferences-cleaned.tabulartext-generation100K<n<1M7 likes160 downloads3y agoHugging Face30it4lia /EMBER_cleaned EMBER Cleaned EMBER Cleaned is a cleaned and AI-ready version of the original EMBER (Endgame Malware Benchmark for Research) dataset, a widely used benchmark for static malware detection on Windows Portable Executable (PE) files. The original EMBER dataset was introduced by Endgame / Elastic as an open benchmark for machine-learning-based malware detection using only static PE-derived features, without executing binaries. This cleaned release preserves that purpose while making the… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/EMBER_cleaned.tabulartabular-classification100K<n<1M0 likes160 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.