datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs_lang_classificationec_classificationarxiv-classificationArxiv Classification: a classification of Arxiv Papers (11 classes).
This dataset is intended for long context classification (documents have all > 4k tokens). Copied from "Long Document Classification From Local Word Glimpses via Recurrent Attention Learning"
@ARTICLE{8675939,
author={He, Jun and Wang, Liqun and Liu, Liu and Feng, Jiao and Wu, Hao},
journal={IEEE Access},
title={Long Document Classification From Local Word Glimpses via Recurrent Attention Learning},
year={2019}… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-classification.Cobot_Magic_classification_of_tableware
Cobot_Magic_classification_of_tableware
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_tableware.task903_deceptive_opinion_spam_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task903_deceptive_opinion_spam_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task903_deceptive_opinion_spam_classification.task902_deceptive_opinion_spam_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task902_deceptive_opinion_spam_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task902_deceptive_opinion_spam_classification.Cobot_Magic_classification_of_fruits_and_vegetables
Cobot_Magic_classification_of_fruits_and_vegetables
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables.AfriMCQA-category-classification
Afri-MCQA cross-modal cultural category classification (MTEB)
Classify the cultural category of an entry from its photograph and the question
about it spoken by a native speaker, across 16 African languages.
Labels index this list:
geography, building, and landmarks
public figure and pop culture
cooking and food
objects, materials, clothing
tranditions, art, and history
brands, products, and companies
plants and animals
people, and everyday life
vehicles and transportation… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-category-classification.Cobot_Magic_classification_of_fruits_and_vegetables_a
Cobot_Magic_classification_of_fruits_and_vegetables_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_classification_of_fruits_and_vegetables_a.mbti_classification_dataset_fullPostspatent-classificationPatent Classification: a classification of Patents and abstracts (9 classes).
This dataset is intended for long context classification (non abstract documents are longer that 512 tokens). Data are sampled from "BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization." by Eva Sharma, Chen Li and Lu Wang
See: https://aclanthology.org/P19-1212.pdf
See: https://evasharma.github.io/bigpatent/
It contains 9 unbalanced classes, 35k Patents and abstracts divided into 3 splits:… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/patent-classification.fr-nfr-classificationdfl_classification_512Vehicle_sounds_classification_datasetfinancial-classification
Dataset Creation
This dataset combines financial phrasebank dataset and a financial text dataset from Kaggle.
Given the financial phrasebank dataset does not have a validation split, I thought this might help to validate finance models and also capture the impact of COVID on financial earnings with the more recent Kaggle dataset.
multilingual-sentiment-classification
MultilingualSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
Sentiment classification dataset with binary
(positive vs negative sentiment) labels. Includes 30 languages and dialects.
Task category
t2c
DomainsReviews, Written
Reference
https://huggingface.co/datasets/mteb/multilingual-sentiment-classification
How to evaluate on this task
You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.multimodal_meme_classification_singapore
Dataset Card for Offensive Memes in Singapore Context
Dataset Details
Dataset Description
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards.
Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.happy-whale-dolphin-classificationZeroshot-Audio-Classification-Instructions
Zeroshot-Audio-Classification-Instructions
Convert audio classification dataset into zero-shot format speech instructions, support both single label and multi-label,
VGGSound
FSD50k
Nonspeech7k
urbansound8K
VocalSound
Emotion
Gender
ESD Emotion
Age
Language
TAU Urban Acoustic Scenes 2022
CochlScene
BirdCLEF_2021
EmoBox
AudioSet
We also converted huge WAV files into MP3 16k sample rate to reduce storage size.To prevent leakage, please do not include test set in training session.… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Zeroshot-Audio-Classification-Instructions.multilingual-scala-classification
ScalaClassification
An MTEB dataset
Massive Text Embedding Benchmark
ScaLa a linguistic acceptability dataset for the mainland Scandinavian languages automatically constructed from dependency annotations in Universal Dependencies Treebanks.
Published as part of 'ScandEval: A Benchmark for Scandinavian Natural Language Processing'
Task category
t2c
Domains
Fiction, News, Non-fiction, Blog, Spoken, Web, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-scala-classification.fiqa-sentiment-classification
Dataset Name
Dataset Description
This dataset is based on the task 1 of the Financial Sentiment Analysis in the Wild (FiQA) challenge. It follows the same settings as described in the paper 'A Baseline for Aspect-Based Sentiment Analysis in Financial Microblogs and News'. The dataset is split into three subsets: train, valid, test with sizes 822, 117, 234 respectively.
Dataset Structure
_id: ID of the data point
sentence: The sentence
target: The target of the… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/fiqa-sentiment-classification.hagrid-classification-512p-dataset
Dataset Card for "hagrid-classification-512p-dataset"
More Information needed
TNews-classification
Dataset Card for "TNews-classification"
More Information needed
code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile.
It is intended to be used for training code natural language classifier.
vqa_plant-disease-classification-merged-datasetfma-genre-classification
FMA Genre Classification Dataset
The FMA Genre Classification Dataset is a subset of the Free Music Archive (FMA), containing audio samples and genre labels for music classification tasks. This version uses the "small" subset of FMA, which contains 8,000 tracks of 30 seconds each, evenly distributed across 8 genres.
Dataset Description
Dataset Summary
This dataset consists of 8,000 audio tracks from the Free Music Archive (FMA), each 30 seconds in length… See the full description on the dataset page: https://huggingface.co/datasets/rpmon/fma-genre-classification.cpc-classification-data
CPC classification datasets
These datasets have been used to train the CPC (Cooperative Patent Classification) classification models mentioned in the article Hähnke, V. D., Wéry, A., Wirth, M., & Klenner-Bajaja, A. (2025). Encoder models at the European Patent Office: Pre-training and use cases. World Patent Information, 81, 102360. https://doi.org/10.1016/j.wpi.2025.102360.
Columns:
publication_number: the patent publication number, the content of the publication can be looked up… See the full description on the dataset page: https://huggingface.co/datasets/mwirth-epo/cpc-classification-data.bi-so101-fruits-classificationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so101_follower",
"total_episodes": 2,
"total_frames": 2910,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/bi-so101-fruits-classification.OnlineShopping-classification
Dataset Card for "OnlineShopping-classification"
More Information needed
profner_classification_master
Binary Classification Dataset: Profession Detection in Tweets
This dataset is a derived version of the original PROFNER task, adapted for binary text classification. The goal is to determine whether a tweet mentions a profession or not.
🧠 Objective
Each example contains:
A tweet_id (document identifier)
A text field (full tweet content)
A label, which has been normalized into two classes:
CON_PROFESION: The tweet contains a reference to a profession.
SIN_PROFESION: The… See the full description on the dataset page: https://huggingface.co/datasets/luisgasco/profner_classification_master.
