datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.arxiv-classifier
arXiv Classifier Data
Usage:
from datasets import load_dataset, DownloadMode
# download from HuggingFace
dataset = load_dataset('mlcore/arxiv-classifier', name=<CONFIG NAME>)
# load from G2
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>)
To force the dataset to be re-generated:
dataset = load_dataset('/share/nikola/arxiv_classifier/data/arxiv-classifier', name=<CONFIG NAME>, download_mode=DownloadMode.FORCE_REDOWNLOAD)
See:… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/arxiv-classifier.LoRA-Samples-Intention-Classifier
Dataset Card for LoRA-Samples-Intention-Classifier
Dataset to fine-tune Qwen3-4B-Instruct-2507-LoRA-Intent-Classifier
Dataset Details
Dataset Description
This dataset includes over 10K samples of prompt-intention id pairs for the AI CS agent generator.It is used to fine-tune a small model that powers this agent, reaching a balance of accuracy, efficiency and cost.
Curated by: Li Tuo
Language(s) (NLP): Chinese (primary), English (partial support)
License:… See the full description on the dataset page: https://huggingface.co/datasets/lituokobe/LoRA-Samples-Intention-Classifier.prompt-safety-classifier-100k
🛡️ Malicious vs. Benign Prompt Classifier & Agent Guardrail Dataset (100,000 Rows)
A 100,000-row multi-class safety dataset engineered specifically for training LLM safety classifiers, guardrail models, prompt injection detectors, and autonomous AI agent tool execution defenses.
📊 Dataset Summary
Unlike standard binary safety datasets, this dataset provides a 6-class safety taxonomy, 1–5 severity scoring, attack technique classification, surface-vs-intent flags… See the full description on the dataset page: https://huggingface.co/datasets/Goutam112/prompt-safety-classifier-100k.insurance-classifier-sft
Insurance Coverage Classifier (Stark Law DHS)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
CPT/HCPCS codes → Stark Law DHS classification + compliance notes
Why download this
Compliance automation for physician self-referral rules. Identify which services are Designated Health Services under Stark Law Section… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/insurance-classifier-sft.high-accuracy-email-classifier
High-Accuracy Email Classification Dataset
Dataset Description
This dataset contains 12,000+ emails across 6 categories, specifically curated for high-accuracy email classification tasks. The dataset achieves 98%+ classification accuracy with appropriate models.
Categories
The dataset includes emails from the following categories:
Category
Count
Description
Emoji
Forum
~2,000
Forum posts, discussions, and community notifications
🗣️
Promotions
~2… See the full description on the dataset page: https://huggingface.co/datasets/jason23322/high-accuracy-email-classifier.math-correctness-classifier_64rollouts
RedaAlami/math-correctness-classifier_64rollouts
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO:… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier_64rollouts.pmjay-classifier-sft
PM-JAY Health Benefit Package Classifier
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical specialty + procedure → PM-JAY HBP code, package name, and rate
Why download this
Automate PM-JAY / Ayushman Bharat claim processing. Map procedures to Health Benefit Packages for pre-authorization and reimbursement.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pmjay-classifier-sft.github_fetch_huggingface_terminal_9127_a9k2d4_derived_emotion_classifier
Emotion Classifier Data
A derived dataset used to train an emotion-classification model.
Description
This dataset combines cleaned and re-labeled samples from upstream sources to produce 8,000 labeled records for emotion classification (anger, joy, sadness, fear, surprise).
Provenance
This dataset was derived from the following source datasets:
TianfuXinqu/github_fetch_huggingface_terminal_9080_m7k2p9_upstream_wiki_factoid… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github_fetch_huggingface_terminal_9127_a9k2d4_derived_emotion_classifier.toxicity-classifier-dataset-v3
Multilingual Toxicity Dataset V3
124K balanced samples for binary toxicity classification across three languages.
Author
Görkem Yıldız
GitHub: gorkem371
Website: gorkemyildiz.com
Dataset Details
Split
Samples
Train
105,945
Valid
18,684
Total
124,629
Languages
Turkish (~34%) — sourced primarily from Overfit-GM
Arabic (~32%) — sourced primarily from arabic-hate-speech-superset
English (~34%) — sourced from toxic_conversations_50k… See the full description on the dataset page: https://huggingface.co/datasets/gorkem371/toxicity-classifier-dataset-v3.VV-classifier-2.0-payment-runs-v1.2refusal-classifier-data
Refusal Classifier Training Data
The english_only subset is taken from the downsampled configuration of agentlans/en-chat-refusal.
The multilingual subset is created by mixing agentlans/multilingual-chat-refusal with the english_only subset in a 1:1 ratio.
All conversations have been formatted using role tokens and ellipses <|...|> for long texts.
The resulting files are shuffled and split 80% training, 20% testing.
VV-classifier-2.0-payment-runs_v1.1math-correctness-classifier-aimes-amo
RedaAlami/math-correctness-classifier-aimes-amo
Dataset Description
This dataset contains mathematical reasoning problems and model responses formatted for training correctness classifiers. Each record includes a problem statement, a model's solution attempt, and a binary label indicating correctness.
The dataset spans three benchmarks:
AIME 2024: American Invitational Mathematics Examination 2024
AIME 2025: American Invitational Mathematics Examination 2025
AMO: Asian… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/math-correctness-classifier-aimes-amo.korean-sign-word-classifier-mediapipe-test-100
KSL Test-100 Evaluation Dataset
This dataset contains 100 isolated Korean sign-language word videos used for the Test-100 evaluation of the MediaPipe keypoint word classifier.
Contents
videos/: MP4 isolated Korean sign-language word clips.
metadata.csv: Video-level labels, model predictions, and evaluation fields.
evaluation/: Full evaluation reports and machine-readable result files.
Evaluation Summary
Model:… See the full description on the dataset page: https://huggingface.co/datasets/Seoyoung07/korean-sign-word-classifier-mediapipe-test-100.context-relevance-classifier-dataset
context-relevance-classifier-dataset
This dataset is designed to train or evaluate models on determining whether an answer to a question is grounded in a given context.
Each sample includes:
question: A question.
answer: A possible answer to the question.
context: A legal passage or reference document.
label:
1 → The answer is supported by the context.
0 → The answer is not supported by the context.
Dataset Source
This dataset is derived from:… See the full description on the dataset page: https://huggingface.co/datasets/axondendriteplus/context-relevance-classifier-dataset.han-humanoid-safety-risk-classifier-v1
Humanoid Safety Risk Classifier (HSRC)
Objective
Predict whether a safety alert should be triggered
based on environmental and actuator conditions.
Problem Type
Binary Classification
Input Features
obstacle_proximity_index
joint_temperature
vibration_level
movement_speed_m_s
torque_index
Output
safety_risk_probability
predicted_label (true/false)
Model Architecture
Feature scaling layer
Dense hidden layers
Sigmoid output… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-humanoid-safety-risk-classifier-v1.cuentas-claras-sat-classifier
Cuentas Claras — SAT Transaction Classifier dataset
Instruction-tuning data that teaches a small model to classify a free-text
transaction description into its SAT account code, deductibility, and IVA
treatment — the core fine-tune (🎯 Well-Tuned) behind the Cuentas Claras accountant
agent.
Format
Chat-format JSONL. Each row:
{
"messages": [
{"role": "system", "content": "Eres un clasificador contable mexicano. ..."},
{"role": "user", "content":… See the full description on the dataset page: https://huggingface.co/datasets/eldinosaur/cuentas-claras-sat-classifier.VV-classifier-2.0-payment-runs-v1.3high-accuracy-email-classifier-indonesian
High-Accuracy Email Classification Dataset Indonesian Translation
This dataset is an Indonesian translation/adaptation of
jason23322/high-accuracy-email-classifier.
Contribution
The original dataset, labels, IDs, and split membership come from
jason23322/high-accuracy-email-classifier, released under the Apache 2.0 license.
This repository contributes the Indonesian translation:
subject, body, and text are translated into Indonesian.
id, category, and category_id… See the full description on the dataset page: https://huggingface.co/datasets/chairulridjal/high-accuracy-email-classifier-indonesian.mockgen-classifier-data-v1
Schema-property semantic classification data (v1)
Field metadata — property name, label, containing entity, neighbouring property names, type, annotations — paired with a semantic hint such as country, currency, person_full_name or gl_account. Each row describes a schema field, never a business record. No values are included.
The task: given only what a schema says about a field, decide what the field means, so a mock-data generator can produce something plausible for it.… See the full description on the dataset page: https://huggingface.co/datasets/Unseen1980/mockgen-classifier-data-v1.contract-classifiergenai-smartcity-classifier
GenAI Smart City Classification Dataset
A curated and augmented dataset for training and evaluating transformer models that classify whether a text (e.g., abstract segment, contribution sentence) describes a Generative AI (GenAI) application in the context of smart cities.
The full codebase for this project can be found [here]here.
1. Dataset Purpose
Supports binary classification:
GenAI used for smart city application
Not related
Used to fine-tune the DeBERTa model in… See the full description on the dataset page: https://huggingface.co/datasets/joaocarlosnb/genai-smartcity-classifier.adaption-legal-clause-classifier-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-legal_clause_classifier
This dataset contains pairs of legal agreement clauses and their corresponding classifications, focusing primarily on termination conditions and intellectual property license grants. Each sample includes a prompt presenting a specific contract excerpt and a completion that identifies the clause type along with relevant entities such as parties, durations, or… See the full description on the dataset page: https://huggingface.co/datasets/avishekh27/adaption-legal-clause-classifier-v1.llama3-8b-instruct-healthcare-dataset-for-classifiersafety_classifier_training_datageo-task-classifierdocument-classifierspider-classifier-training-data
Spider Classifier — Training Manifest
Public release of the training manifest used to fine-tune the
Spiders of New Hampshire species
classifier.
This manifest enumerates every photo used to train, validate, and test the
model. Each row links back to the original observation and photo on
iNaturalist, preserving full attribution and
license metadata.
Source
Model run: 20260528_104624_licensed_dinov2_l_14_reg4_518
Generated: 2026-05-29T02:17:00.691986+00:00… See the full description on the dataset page: https://huggingface.co/datasets/bwirth/spider-classifier-training-data.imdb-sentiment-classifier-evals
