datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
student-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
ni-20-categoriesreachy-mini-app-categoriesvehicle-inspection-defect-categories
Vehicle inspection defect categories: MOT major, minor and dangerous verdicts
Canonical, always-current version: https://referencesource.org/vehicle-inspection-defect-categories/
Machine-readable: https://referencesource.org/vehicle-inspection-defect-categories/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-06
Stale after: 2027-02-02 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 1422
What each… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/vehicle-inspection-defect-categories.model-categories
Model Categories
31 working models organized by use case.
🚀 dispatchAI
agda-categories-informalized
agda-categories, informalized
4,541 declarations from the agda-categories
library, each paired with an informal, LaTeX-flavoured natural-language
statement written by GLM-5.2. The natural language is written to be precise
enough to re-formalise from, so the intended use is training a model to
reconstruct the formal Agda source from the prose alone.
Declarations were extracted with a fork of Agda
that dumps one JSON record per named declaration (with its full source range)
during… See the full description on the dataset page: https://huggingface.co/datasets/astral-expmath/agda-categories-informalized.categories-100-en
Categories-100-en
A curated set of 100 human-readable English category names, generated via self-instruct prompting with DeepSeek V4 Flash 0731.
Use cases
As a label set for zero-shot or few-shot text classification demos
For testing long-list selection / tag extraction in LLMs
As a quick vocabulary of common subject areas for educational tools
Limitations
Categories are flat (no hierarchy) and some overlap (e.g., "Physics" vs."Astronomy", "Art"… See the full description on the dataset page: https://huggingface.co/datasets/RowRed/categories-100-en.BEA2019_categoriescoco-panoptic-categorieslightspeed-categories
Lightspeed Categories
This is a database of the categories of the Lightspeed Filter. You can use this with the Lightspeed API, which will provide you the category number for a URL.
SuperBEIR-categories-with-rationales-gflSemEval_categoriesSwag_categories_doubledstudent-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset.
It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics.
SemEval_categories_doubledaftermath_questions_to_categories
Aftermath of DrawEduMath
This contains questions_to_categories.json, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains a mapping from DrawEduMath questions to question types, following the taxonomy introduced in the original DrawEduMath paper.
Please consult the datacard for DrawEduMath for detailed information about data source.
Quick links:… See the full description on the dataset page: https://huggingface.co/datasets/lucy3/aftermath_questions_to_categories.adaption-rupeeroast-txn-categories-v1
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-rupeeroast_txn_categories
This dataset contains categorized Indian UPI and bank transaction records, mapping merchant names to spending categories such as Food, Investment, and Transfers. It features multilingual support with entries in English, Hindi, and Hinglish across 13 distinct financial categories. The data is designed for training financial AI models to… See the full description on the dataset page: https://huggingface.co/datasets/shreya-var/adaption-rupeeroast-txn-categories-v1.mercari-jmty-categories-serveycarbon-credit-categories
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
carbon_credit_categories
This dataset consists of pairs mapping unique carbon credit project identifiers to their specific mitigation categories, such as Direct Air Capture (DAC), Renewable Energy, and Afforestation. The content appears structured for classification or lookup tasks within the voluntary carbon market domain. Each entry links a standardized ID code to a concise label representing… See the full description on the dataset page: https://huggingface.co/datasets/joduor/carbon-credit-categories.BEA2019_categories_doubledGPC-categorieslogolens-categories
Dataset Name: LogoLens Categories
Description
The logolens-categories dataset contains classifications of logo designs based on type, style attributes, and color properties. It is ideal for tasks related to visual analysis, design inspiration, and automated logo classification.
Version: 1.0.0
Homepage: Hugging Face Dataset
License: MIT
Features
Field
Type
Description
Example
category_type
string
The classification group (e.g., logo_types).… See the full description on the dataset page: https://huggingface.co/datasets/Tiny-Factories/logolens-categories.books-ner-dataset-categories
Books Named Entity Recognition (NER) Dataset
A small Named‑Entity‑Recognition (NER) corpus built from categories contained in Project Gutenberg’s public catalogues. It is intended for training or benchmarking entity extractors such as GLiNER on bibliographic metadata.
1 Provenance
This dataset provenance originates from Project Gutenberg's public catalogue.
2 Quick facts
Records (total)
2364
Train split
2127 queries
Eval split
237 queries… See the full description on the dataset page: https://huggingface.co/datasets/empathyai/books-ner-dataset-categories.Swag_categoriesmecari-categories-survayadaption-rupeeroast-txn-categories
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-rupeeroast_txn_categories
This dataset contains categorized Indian UPI and bank transaction records, mapping merchant names to spending categories such as Food, Investment, and Transfers. It features multilingual support with entries in English, Hindi, and Hinglish across 13 distinct financial categories. The data is designed for training financial AI models to… See the full description on the dataset page: https://huggingface.co/datasets/shreya-var/adaption-rupeeroast-txn-categories.
