CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.9k downloads2y agoHugging Face02MicPie /unpredictable_cluster00The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choicen<1K0 likes152 downloads4y agoHugging Face03Malikeh1375 /clustered_tulu_3_16 Clustered_Tulu_3_16 Multi-Domain Dataset This dataset contains high-quality examples across 16 specialized domains, automatically extracted and curated from the Tulu-3 SFT mixture using advanced clustering techniques. 🎯 Multi-Domain Structure This repository provides 16 domain-specific configurations, each optimized for different types of tasks: Configuration Domain Train Test Total python_string_and_list_processing Python String & List Processing 43,564 10… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/clustered_tulu_3_16.texttext-generation100K<n<1M0 likes124 downloads1y agoHugging Face04argilla /FinePersonas-v0.1-clustering-100k Dataset Card for PersonaHub FineWeb-Edu 4 Clustering 100k This dataset has been created with distilabel. The following figure is a map of the clusters generated from the pipeline. It's automatically generated by the TextClustering with all the information gathered. It contains 177 different clusters, which were assigned a set of 3 labels each, and the black dots correspond to those unclassified examples. Dataset Summary This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-v0.1-clustering-100k.texttext-generation100K<n<1M13 likes80 downloads2y agoHugging Face05Galtea-AI /galtea-red-teaming-clustered-data Galtea Red Teaming: Non-Commercial Subset This dataset contains a curated collection of adversarial prompts used for red teaming and LLM safety evaluation. All prompts come from datasets under non-commercial licenses and have been: Deduplicated Normalized into a consistent format Automatically clustered based on semantic meaning Each entry includes: prompt: the adversarial instruction source: the dataset of origin cluster: a numeric cluster ID based on prompt behavior… See the full description on the dataset page: https://huggingface.co/datasets/Galtea-AI/galtea-red-teaming-clustered-data.texttext-generation10K<n<100K1 likes68 downloads1y agoHugging Face06MicPie /unpredictable_cluster08The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choicen<1K0 likes57 downloads4y agoHugging Face07ymoslem /AIME-24-25-26-clustered AIME 2024-2026, clustered 90 AIME problems from 2024, 2025 and 2026, embedded and assigned to the three semantic clusters used by CRE-Router. Intended as a held-out test set for query routing: three years of 30 problems each, spread evenly across clusters, so accuracy can be broken down by year. Figure rule Every problem is self-contained. Where a figure is needed to solve a problem it is included as source, Asymptote or plain text; where a figure was decorative… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/AIME-24-25-26-clustered.tabulartext-generationn<1K0 likes56 downloads19d agoHugging Face08MicPie /unpredictable_cluster20The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes54 downloads4y agoHugging Face09Malikeh1375 /clustered_tulu_3_8 Clustered_Tulu_3_8 Multi-Domain Dataset This dataset contains high-quality examples across 8 specialized domains, automatically extracted and curated from the Tulu-3 SFT mixture using advanced clustering techniques. 🎯 Multi-Domain Structure This repository provides 8 domain-specific configurations, each optimized for different types of tasks: Configuration Domain Train Test Total programming_and_code_development Programming & Code Development 88,783 22,196 110… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/clustered_tulu_3_8.texttext-generation100K<n<1M1 likes54 downloads1y agoHugging Face10MicPie /unpredictable_cluster-noiseThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes46 downloads4y agoHugging Face11MicPie /unpredictable_cluster15The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes40 downloads4y agoHugging Face12MicPie /unpredictable_cluster29The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes39 downloads4y agoHugging Face13MicPie /unpredictable_cluster07The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes38 downloads4y agoHugging Face14MicPie /unpredictable_cluster06The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choicen<1K0 likes37 downloads4y agoHugging Face15MicPie /unpredictable_cluster26The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes35 downloads4y agoHugging Face16MicPie /unpredictable_cluster28The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes33 downloads4y agoHugging Face17MicPie /unpredictable_cluster09The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes33 downloads4y agoHugging Face18MicPie /unpredictable_cluster14The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes32 downloads4y agoHugging Face19MicPie /unpredictable_cluster27The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes30 downloads4y agoHugging Face20MicPie /unpredictable_cluster12The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes28 downloads4y agoHugging Face21MicPie /unpredictable_cluster18The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes28 downloads4y agoHugging Face22Ibisbill /Clustering_deduplicated_reasoning Clustering_deduplicated_reasoning 数据集描述 Clustering deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category 文件结构 clustering_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式) 数据格式 数据集包含以下字段: question: str quality: int difficulty: int topic: str validity: int 使用方法 方法1: 使用datasets库 from datasets import load_dataset # 加载数据集 dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Clustering_deduplicated_reasoning.texttext-generation10K<n<100K0 likes28 downloads1y agoHugging Face23MicPie /unpredictable_cluster21The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes27 downloads4y agoHugging Face24MicPie /unpredictable_cluster24The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes27 downloads4y agoHugging Face25MicPie /unpredictable_cluster02The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes26 downloads4y agoHugging Face26MicPie /unpredictable_cluster16The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes25 downloads4y agoHugging Face27MicPie /unpredictable_cluster01The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes24 downloads4y agoHugging Face28MicPie /unpredictable_cluster10The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice1K<n<10K0 likes23 downloads4y agoHugging Face29MicPie /unpredictable_cluster05The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choicen<1K0 likes23 downloads4y agoHugging Face30TatarNLPWorld /tatar-news-clustergated Dataset Card for Tatar News Clustered Dataset Dataset Details Dataset Description The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.texttext-classification10K<n<100K0 likes23 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.