datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.unpredictable_cluster00The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.clustered_tulu_3_16
Clustered_Tulu_3_16 Multi-Domain Dataset
This dataset contains high-quality examples across 16 specialized domains, automatically extracted and curated from the Tulu-3 SFT mixture using advanced clustering techniques.
🎯 Multi-Domain Structure
This repository provides 16 domain-specific configurations, each optimized for different types of tasks:
Configuration
Domain
Train
Test
Total
python_string_and_list_processing
Python String & List Processing
43,564
10… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/clustered_tulu_3_16.FinePersonas-v0.1-clustering-100k
Dataset Card for PersonaHub FineWeb-Edu 4 Clustering 100k
This dataset has been created with distilabel.
The following figure is a map of the clusters generated from the pipeline. It's automatically generated by the TextClustering with all the information gathered.
It contains 177 different clusters, which were assigned a set of 3 labels each, and the black dots correspond to those unclassified examples.
Dataset Summary
This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-v0.1-clustering-100k.galtea-red-teaming-clustered-data
Galtea Red Teaming: Non-Commercial Subset
This dataset contains a curated collection of adversarial prompts used for red teaming and LLM safety evaluation. All prompts come from datasets under non-commercial licenses and have been:
Deduplicated
Normalized into a consistent format
Automatically clustered based on semantic meaning
Each entry includes:
prompt: the adversarial instruction
source: the dataset of origin
cluster: a numeric cluster ID based on prompt behavior… See the full description on the dataset page: https://huggingface.co/datasets/Galtea-AI/galtea-red-teaming-clustered-data.unpredictable_cluster08The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.AIME-24-25-26-clustered
AIME 2024-2026, clustered
90 AIME problems from 2024, 2025 and 2026, embedded and assigned to the three
semantic clusters used by CRE-Router.
Intended as a held-out test set for query routing: three years of 30 problems
each, spread evenly across clusters, so accuracy can be broken down by year.
Figure rule
Every problem is self-contained. Where a figure is needed to solve a problem it
is included as source, Asymptote or plain text; where a figure was decorative… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/AIME-24-25-26-clustered.unpredictable_cluster20The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.clustered_tulu_3_8
Clustered_Tulu_3_8 Multi-Domain Dataset
This dataset contains high-quality examples across 8 specialized domains, automatically extracted and curated from the Tulu-3 SFT mixture using advanced clustering techniques.
🎯 Multi-Domain Structure
This repository provides 8 domain-specific configurations, each optimized for different types of tasks:
Configuration
Domain
Train
Test
Total
programming_and_code_development
Programming & Code Development
88,783
22,196
110… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/clustered_tulu_3_8.unpredictable_cluster-noiseThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster15The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster29The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster07The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster06The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster26The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster28The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster09The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster14The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster27The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster12The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster18The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Clustering_deduplicated_reasoning
Clustering_deduplicated_reasoning
数据集描述
Clustering deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
clustering_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
# 加载数据集
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Clustering_deduplicated_reasoning.unpredictable_cluster21The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster24The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster02The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster16The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster01The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster10The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unpredictable_cluster05The UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.tatar-news-cluster
Dataset Card for Tatar News Clustered Dataset
Dataset Details
Dataset Description
The Tatar News Clustered Dataset is a comprehensive collection of 57,340 Tatar language news articles with topic categories, curated by TatarNLPWorld as part of the Tat2Vec project. The dataset includes full article content, titles, source names, publication dates, and 282 topic categories. It is designed for multi-class text classification, topic modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-cluster.
