datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-bank-transaction-categories-v2
US Bank Transaction Categories v2 — Synthetic Dataset
68,000 sign-prefixed transaction descriptions across 17 spending categories, modeled after real US bank statement formats. Designed for training classifiers that work on actual bank data — not the clean "Starbucks coffee" descriptions that most datasets use.
Successor to v1.
Why This Dataset
Real bank transaction data is private. But the formats are universal — Chase, Apple Card, PayPal, Capital One, Mercury all… See the full description on the dataset page: https://huggingface.co/datasets/DoDataThings/us-bank-transaction-categories-v2.Yahoo_Answers_10_categories_for_NLP
Dataset Card for Dataset Name
The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information.
Dataset Description
The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.Politifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
libero-plus-episode-categories
LIBERO-Plus episode → perturbation-category labels
Which perturbation category each of the 14,347 LIBERO-Plus episodes belongs to.
LIBERO-Plus perturbs a base LIBERO task along several axes. The release these
labels came from contains five of them --- the project describes more, so treat
this as the taxonomy of this 14,347-episode release, not of LIBERO-Plus as a
whole. The
lerobot/libero_plus conversion does not carry that label, so you cannot ask
"how does my policy do under… See the full description on the dataset page: https://huggingface.co/datasets/katsukiono/libero-plus-episode-categories.313k-prices-nyc-vs-la-57-categories
313,126 prices: 57 categories, 2 U.S. ZIPs, 29 days
NYC vs LA Retail Prices Raw Dataset (2026)
How do listed and package-standardized prices vary between selected New York and Los Angeles ZIP markets across 57 everyday categories and 29 days?
This fixed research snapshot contains 313,126 unaggregated, quality-filtered price observations across 57 categories, 2 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. The analysis-ready CSV preserves… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/313k-prices-nyc-vs-la-57-categories.us-bank-transaction-categories
Update Notice
A newer version of this dataset is available: DoDataThings/us-bank-transaction-categories-v2
v2 adds [debit]/[credit] sign prefixes, 24,000 samples (up from 16,000), 500+ merchants, PayPal wrapper patterns across all categories, and a refined 16-category taxonomy (Housing split into Rent; Mortgage removed). See the v2 dataset card for details.
This v1 dataset remains available. If you don't need sign-aware data, v1 is still useful for basic transaction classification… See the full description on the dataset page: https://huggingface.co/datasets/DoDataThings/us-bank-transaction-categories.us-bank-transaction-categories
Update Notice
A newer version of this dataset is available: DoDataThings/us-bank-transaction-categories-v2
v2 adds [debit]/[credit] sign prefixes, 24,000 samples (up from 16,000), 500+ merchants, PayPal wrapper patterns across all categories, and a refined 16-category taxonomy (Housing split into Rent; Mortgage removed). See the v2 dataset card for details.
This v1 dataset remains available. If you don't need sign-aware data, v1 is still useful for basic transaction… See the full description on the dataset page: https://huggingface.co/datasets/DEVILHADYOURMOM/us-bank-transaction-categories.news-categories
English News Headline Dataset
Overview
This dataset contains 50,000 English news headlines categorized into 10 topical classes, designed for text classification and NLP studies such as news topic modeling, transfer learning, and zero‑shot evaluation.
Each record includes:
title: news headline text
topic: one of ten predefined categories
genre: one of four predefined descriptor of the story style (e.g., Informational, Analysis)
source: media outlet name
date:… See the full description on the dataset page: https://huggingface.co/datasets/momentum-lab/news-categories.news-categoriesNSINA-Categories
Sinhala News Category Prediction
This is a text classification task created with the NSINA dataset. This dataset is also released with the same license as NSINA.
Data
Data can be loaded into pandas dataframes using the following code.
from datasets import Dataset
from datasets import load_dataset
train = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='train'))
test = Dataset.to_pandas(load_dataset('sinhala-nlp/NSINA-Categories', split='test'))… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/NSINA-Categories.amazon-all-categories-best-sellers-reviewswhat-things-cost-57-categories-12-us-zips
1,861,051 prices: 57 categories, 12 U.S. ZIPs, 29 days
1.86M Prices: What Things Cost in 12 U.S. ZIPs
What does a hamburger bun, a bag of dog food, a box of tampons, an air conditioner, or an NVMe SSD cost in different U.S. ZIP markets?
This fixed research snapshot contains 1,861,051 unaggregated, quality-filtered retail price observations across 57 product categories, 12 U.S. ZIP markets, and 29 consecutive dates from July 21 through August 18, 2026. One analysis-ready CSV… See the full description on the dataset page: https://huggingface.co/datasets/costinflation/what-things-cost-57-categories-12-us-zips.turkish-news-categorieswikipedia_categoriesBanglaSumXL_CategoriesWe upgraded our dataset with category. Another column category is added to the dataset which tells which category the text belongs to. We picked 7 categories and manually annotated the dataset over 7 categories, International, State, Entertainment, Economy, Education, and Technology. This dataset will be useful for NLP tasks like Bangla text classification. There are rarely any proper text classification datasets in low-resource languages like Bangla. There are some sentiment classification… See the full description on the dataset page: https://huggingface.co/datasets/midnightGlow/BanglaSumXL_Categories.ai-categories
Bilarna Service Taxonomy (41 categories)
Dataset Summary
This dataset contains a compact taxonomy of service/category labels associated with a single website (https://bilarna.com). It is designed for lightweight experiments in category normalization, labeling, and taxonomy alignment (e.g., mapping internal product/service categories to a standardized catalog).
Each row includes:
id: numeric identifier
website: source website (constant in this dataset)
name:… See the full description on the dataset page: https://huggingface.co/datasets/erhankocabas/ai-categories.wolo-app-categories-to-descriptionwikipedia_categories_labelssupermarket_categoriescensus-job-categories
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/andrew-behr/census-job-categories.all_4_categories_finetuningCategories-1k-Globalyproducts_categories_dataecom_categoriesautotrain-data-question-categoriescategories-11kcategories-1kai_tweet_categories10k tweets labeled in a Multilabel fashion for 8 categories :
'AI News', 'AI Tools', 'AI Research', 'AI Models', 'AI Usecases','AI Open Source', 'Podcasts/Talks/Events', 'AI Opinions','Non AI'
questionsdata_with_categoriesconcerning_interactions_with_categories
