datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BRIGHTER-emotion-categories
BRIGHTER Emotion Categories Dataset
This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages.
Dataset Description
The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.wikipedia-categories
WikipediaPTCategoriesClusteringP2P
Cluster the first paragraph of Brazilian Portuguese Wikipedia articles into 15 broad subject categories (História, Geografia, Política, Esporte, Música, Cinema, Literatura, Religião, Ciência, Tecnologia, Animais, Plantas, Medicina, Filosofia, Astronomia). Articles sampled via MediaWiki API category traversal (depth 2) of curated root categories; first paragraph from the extracts API.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/wikipedia-categories.maltese_news_categories
Maltese News Categories
A multi-label topic classification dataset for Maltese News Articles.
Data Collection
The data was collected from the press_mt subset from Korpus Malti v4.0.
Article contents were cleaned to filter out JavaScript, CSS, & repeated non-Maltese sub-headings.
The labels are based on the category field from this corpus.
Additional filtering & cleaning was performed as follows:
Documents with generic categories (News, Local, Headlines, Uncategorised… See the full description on the dataset page: https://huggingface.co/datasets/MLRS/maltese_news_categories.flowers-102-categoriesstudent-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
ni-20-categoriesus-bank-transaction-categories-v2
US Bank Transaction Categories v2 — Synthetic Dataset
68,000 sign-prefixed transaction descriptions across 17 spending categories, modeled after real US bank statement formats. Designed for training classifiers that work on actual bank data — not the clean "Starbucks coffee" descriptions that most datasets use.
Successor to v1.
Why This Dataset
Real bank transaction data is private. But the formats are universal — Chase, Apple Card, PayPal, Capital One, Mercury all… See the full description on the dataset page: https://huggingface.co/datasets/DoDataThings/us-bank-transaction-categories-v2.zlib-categoriesshire-vehicle-categories
🤖 ShIRE Dataset: 6D Pose Trajectories of Independently Moving Objects
The ShIRE Dataset is the official reference implementation of the method introduced in the paper:
ShIRE: Real-Time 6D Pose Tracking of Independently Moving Objects
Presented at VISAPP 2021
📝 Overview
ShIRE (Short for Shape-aware Independent Real-time Estimation) is a method for real-time detection and 6D pose trajectory estimation of independently moving objects in dynamic environments.
This… See the full description on the dataset page: https://huggingface.co/datasets/langutang/shire-vehicle-categories.vault-categoriesthe-eye-categoriesreachy-mini-app-categoriesecommerce-product-classification-by-categories
🛍️ E-Commerce Product Classification Dataset
Dataset Description
This dataset contains 31,851 realistic English product descriptions across 33 e-commerce categories, designed for training product classification models.
Key Features
✅ 31,851 samples with natural language variations
✅ 33 product categories covering major e-commerce segments
✅ 4 seller personas: Professional, Individual, Reseller, Minimal
✅ Realistic noise: Typos, abbreviations, casual language… See the full description on the dataset page: https://huggingface.co/datasets/Lezh1n/ecommerce-product-classification-by-categories.wikipedia-categories-largefashionpedia_4_categories
Dataset Card for Fashionpedia_4_categories
This dataset is a variation of the fashionpedia dataset available here, with 2 key differences:
It contains only 4 categories:
Clothing
Shoes
Bags
Accessories
New splits were created:
Train: 90% of the images
Val: 5%
Test 5%
The goal is to make the detection task easier with 4 categories instead of 46 for the full fashionpedia dataset.
This dataset was created using the detection_datasets library (GitHub, PyPI), you can check here the… See the full description on the dataset page: https://huggingface.co/datasets/detection-datasets/fashionpedia_4_categories.CLEVR_categoriesmmlu_pro_categories
MMLU-Pro Dataset : Per-Category Splits
This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories.
from datasets import load_dataset
ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology')
The available tasks are:
Category Name
Split Name
Biology
category_biology
Business
category_business
Chemistry
category_chemistry
Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.arxiv_categories📄 Paper: Efficient Few-shot Learning for Multi-label Classification of Scientific Documents with Many Classes (ICNLSP 2024)
💻 GitHub: https://github.com/sebischair/FusionSent
This is a dataset of scientific documents derived from arXiv metadata. The arXiv metadata provides information about more than 2 million scholarly articles published in arXiv from various scientific fields. We use this metadata to create a dataset of 203,961 titles and abstracts categorized into 130 different classes.… See the full description on the dataset page: https://huggingface.co/datasets/TimSchopf/arxiv_categories.wikidata-enwiki-categories-and-statementsYahoo_Answers_10_categories_for_NLP
Dataset Card for Dataset Name
The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information.
Dataset Description
The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.ia-books-categorieslibgen-categoriesGLDv2_Top_51_Categories
Dataset Card for Dataset Name
Dataset Summary
This dataset is a subset of Kaggle's Google Landmark Recognition 2021 competition with only the categories with more than 500 images.
https://www.kaggle.com/competitions/landmark-recognition-2021/data
The dataset consists of a total of 45579 224x224 color images in 51 categories.
Languages
English
Dataset Structure
Data Fields
landmark_id: Int - Numeric identifier of the category
category :… See the full description on the dataset page: https://huggingface.co/datasets/pemujo/GLDv2_Top_51_Categories.wikipedia-categoriesThis dataset is derived from the research titled 'WC-SBERT: Zero-Shot Topic Classification Using SBERT and Self-Training with Wikipedia Categories.'
It is based on the Hugging Face Wikipedia 20220301.en dataset.
The 'id' represents the Wikipedia page ID, and the 'categories' are the categories associated with each page.
emotion-analysis-8-categories-datasetPolitifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
reddit_categories_clean
Dataset Card for "reddit_categories_clean"
More Information needed
libero-plus-episode-categories
LIBERO-Plus episode → perturbation-category labels
Which perturbation category each of the 14,347 LIBERO-Plus episodes belongs to.
LIBERO-Plus perturbs a base LIBERO task along several axes. The release these
labels came from contains five of them --- the project describes more, so treat
this as the taxonomy of this 14,347-episode release, not of LIBERO-Plus as a
whole. The
lerobot/libero_plus conversion does not carry that label, so you cannot ask
"how does my policy do under… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/libero-plus-episode-categories.dawiki_categories
Danish Wikipedia Categories
The dataset was created entirely from the last Danish Wikipedia dump
by traversing the category hierarchy in the categorylinks table.
All categories that were one level bellow the topcategories, and which had more than 30 articles assigned to them were selected.
In order to see whether an article belongs to a certain category I checked, whether the article was connected to the category in the directed graph of the category hierarchy.
If the length of the… See the full description on the dataset page: https://huggingface.co/datasets/kardosdrur/dawiki_categories.313k-prices-nyc-vs-la-57-categories
313,126 prices: 57 categories, 2 U.S. ZIPs, 29 days
NYC vs LA Retail Prices Raw Dataset (2026)
How do listed and package-standardized prices vary between selected New York and Los Angeles ZIP markets across 57 everyday categories and 29 days?
This fixed research snapshot contains 313,126 unaggregated, quality-filtered price observations across 57 categories, 2 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready CSV preserves… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/313k-prices-nyc-vs-la-57-categories.
