datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs_lang_classificationjigsaw-toxic-comment-classification-challenge
Dataset Description
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
You must create a model which predicts a probability of each type of toxicity for each comment.
File descriptions
train.csv - the training set, contains comments with their binary labels
test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.Barkopedia_Dog_Sex_Classification_Dataset
📦 Dataset Description
This dataset is part of the Barkopedia Challenge: https://uta-acl2.github.io/barkopedia.html
Check training data on Hugging Face:
👉 ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
29,345 total clips of vocalizations from 156 individual dogs across 5 breeds:
Shiba Inu
Husky
Chihuahua
German Shepherd
Pitbull
Training set: 26,895 clips
13,567 female13,328 male
Test set: 2,450… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset.ec_classificationaidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.eai-taxonomy-math-w-fm-classify-behaviors
🧮 EAI Taxonomy Math w/ Behavioral Classifications (10K Sample)
A 10,000 document sample from EssentialAI/eai-taxonomy-math-w-fm enhanced with 4 behavioral reasoning classifications using GPT-4.1-mini.
Behavioral Classifications
Structured behavioral analysis following the approach from cognitive-behaviors:
backtracking_json: Identifies reasoning that backtracks or revisits earlier steps
backward_chaining_json: Detects goal-oriented reasoning working backwards… See the full description on the dataset page: https://huggingface.co/datasets/nlile/eai-taxonomy-math-w-fm-classify-behaviors.jailbreak-classification
Jailbreak Classification
Dataset Summary
Dataset used to classify prompts as jailbreak vs. benign.
Dataset Structure
Data Fields
prompt: an LLM prompt
type: classification label, either jailbreak or benign
Dataset Creation
Curation Rationale
Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs.
Source Data
Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.classified_images_gemmaopenclaw-classification-dataset
OpenClaw GitHub Interest Classification Dataset
This folder is a small, maintainable dataset for improving OpenClaw GitHub PR and
issue classification. It is intentionally separate from the notifier runtime so
it can be edited locally, reviewed in source control, or uploaded as a Hugging
Face dataset repository.
Canonical Hugging Face dataset: dutifuldev/openclaw-classification-dataset
URL: https://huggingface.co/datasets/dutifuldev/openclaw-classification-dataset
The current… See the full description on the dataset page: https://huggingface.co/datasets/dutifuldev/openclaw-classification-dataset.classnodaikirainajoshitokekkonsurukotoninatta
Bangumi Image Base of Class No Daikirai Na Joshi To Kekkon Suru Koto Ni Natta.
This is the image base of bangumi Class no Daikirai na Joshi to Kekkon suru Koto ni Natta., we detected 33 characters, 3154 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/classnodaikirainajoshitokekkonsurukotoninatta.ILUR-news-text-classification-corpus
News Texts Dataset
We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens).
For more details, refer to the paper.
ClassEval
Dataset Card for FudanSELab ClassEval
Dataset Summary
We manually build ClassEval of 100 class-level Python coding tasks, consists of 100 classes and 412 methods, and average 33.1 test cases per class.
For 100 class-level tasks, diversity is maintained by encompassing these tasks over a wide spectrum of topics, including Management Systems, Data Formatting, Mathematical Operations, Game Development, File Handing, Database Operations and Natural Language Processing.
For… See the full description on the dataset page: https://huggingface.co/datasets/FudanSELab/ClassEval.cpc-classificationsclassimgman2000classimgman2000-2classicstars
Bangumi Image Base of Classic★stars
This is the image base of bangumi Classic★Stars, we detected 68 characters, 6165 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/classicstars.cybersecurity-classification-benchmark
TorchSight Cybersecurity Classification Benchmark
A two-tier benchmark dataset for evaluating cybersecurity document
classifiers, released with the TorchSight system. Used in:
Dobrovolskyi, I. Security Document Classification with a Fine-Tuned Local
Large Language Model: Benchmark Data and an Open-Source System. Journal of
Information Security and Applications, 2026.
Canonical per-model numbers live in BENCHMARK_NUMBERS.md,
auto-generated from the per-prediction result JSONs… See the full description on the dataset page: https://huggingface.co/datasets/torchsight/cybersecurity-classification-benchmark.ParlaSpeech-RS
The Serbian Parliamentary Spoken Dataset ParlaSpeech-RS 1.0
The master dataset can be found at http://hdl.handle.net/11356/1834.
Notice: ParlaSpeech corpora are currently in the process of enrichment with new features. Follow our progress here: http://clarinsi.github.io/parlaspeech
The ParlaSpeech-RS dataset is built from the transcripts of parliamentary proceedings available in the Serbian part of the ParlaMint corpus (http://hdl.handle.net/11356/1859), and the parliamentary… See the full description on the dataset page: https://huggingface.co/datasets/classla/ParlaSpeech-RS.ade_corpus_v2_classification
ADE-Corpus-V2 Dataset: Adverse Drug Reaction Data.
This is a dataset for classification if a sentence is ADE-related (True) or not (False).
Train size: 17,637
Test size: 5,879
Source dataset
Paper
image-for-LULU-classifier
LULU Generated Images (SD 1.4)
Synthetic images for 79 COCO object classes, generated with Stable Diffusion v1.4.
Layout (original paths)
explicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
implicit/{category}/sample_{idx:05d}_p{prompt}_s{seed}.png
explicit: 79 classes × 30 prompts × 100 seeds = 237,000 images (~98 GB)
implicit: validation-style prompts, fewer samples per class
Filename fields:
prompt_idx: 0–29
seed: 0–99 (explicit)… See the full description on the dataset page: https://huggingface.co/datasets/weilulobster/image-for-LULU-classifier.sd-webui-forge-classicStable Diffusion WebUI Forge - Classic
[ Classic | Neo ]
Stable Diffusion WebUI Forge is a platform on top of the original Stable Diffusion WebUI by AUTOMATIC1111, to make development easier, optimize resource management, speed up inference, and study experimental features.
The name "Forge" is inspired by "Minecraft Forge". This project aims to become the Forge of Stable Diffusion WebUI.
- lllyasviel
(paraphrased)
"Classic" mainly serves as an archive for the "previous" version of… See the full description on the dataset page: https://huggingface.co/datasets/WhiteAiZ/sd-webui-forge-classic.butteflies_with_classesCourtesy: https://www.kaggle.com/datasets/phucthaiv02/butterfly-image-classification
arxiv-classificationArxiv Classification: a classification of Arxiv Papers (11 classes).
This dataset is intended for long context classification (documents have all > 4k tokens). Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning"
@ARTICLE{8675939,
author={He, Jun and Wang, Liqun and Liu, Liu and Feng, Jiao and Wu, Hao},
journal={IEEE Access},
title={Long Document Classification From Local Word Glimpses via Recurrent Attention Learning},
year={2019}… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-classification.commit-classification-dataset
Commit Classification Dataset
This dataset is designed for multi-label classification of Git commit messages into predefined categories.
Dataset Summary
This dataset contains:
Training data: Commit messages and their corresponding labels for training the model.
Validation data: A separate set of messages for tuning and evaluation.
Testing data: Unlabeled commit messages for testing the model’s performance.
The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.Cobot_Magic_classification_of_tableware
Cobot_Magic_classification_of_tableware
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.data-csgo-weapon-classification
Dataset for project: csgo-weapon-classification
Dataset Description
This dataset has for project csgo-weapon-classification was collected with the help of a bulk google image downloader.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1768x718 RGB PIL image>",
"target": 0
},
{
"image": "<716x375 RGBA PIL image>"… See the full description on the dataset page: https://huggingface.co/datasets/Kaludi/data-csgo-weapon-classification.task903_deceptive_opinion_spam_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.molmo2-tulu4-classified
