datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
Multi-Lingual-Lyrics-for-Genre-Classificationfrom https://www.kaggle.com/datasets/mateibejan/multilingual-lyrics-for-genre-classification
multiclasstask1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.functional-multiclass-gamba
GAMBA Functional Region Multiclass
This representation benchmark asks whether frozen sequence embeddings
separate genomic functional categories. Each row is one annotated region;
label == category.
Loading
from datasets import load_dataset
full_bidi = load_dataset(
"Taykhoom/functional-multiclass-gamba",
"full-bidi",
split="all",
)
paper_test = full_bidi.filter(
lambda row: row["split"] == "test"
and row["category"] != "noncoding_regions"
)… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/functional-multiclass-gamba.DeBERTa_multi-class_cb_datasettask638_multi_woz_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task638_multi_woz_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task638_multi_woz_classification.synthetic-text-classification-news-multi-label
Dataset Card for synthetic-text-classification-news-multi-label
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.toxicity-multi-label-classifier
Part of a course titled "Generative AI application design & development"
https://genai.acloudfan.com/
Created from a dataset available on Kaggle.
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data
casa_3sec_multiclassreddit-AITA-submissions-and-comments-multiclassmulti_edu_classificationmulti-class_cyberbullying_emoji_datasethuman_multi_classifications_500YOLOv8-Multiclass-Object-Detection-Dataset
DATASET SAMPLE
Duality.ai just released a 1000 image dataset used to train a YOLOv8 model in multiclass object detection -- and it's 100% free!
Just create an EDU account here.
This HuggingFace dataset is a 20 image and label sample, but you can get the rest at no cost by creating a FalconCloud account. Once you verify your email, the link will redirect you to the dataset page.
What makes this dataset unique, useful, and capable of bridging the Sim2Real gap?
The digital twins are… See the full description on the dataset page: https://huggingface.co/datasets/duality-robotics/YOLOv8-Multiclass-Object-Detection-Dataset.cicflow-ids-multiclass
CICFlow Multiclass Intrusion Detection Dataset
Overview
This dataset provides a multiclass network intrusion detection (IDS) benchmark derived from CICFlowMeter flow-level features.
It is designed for attack-type classification, robust IDS research, and interpretable security modeling.
Each network flow is labeled as either benign or one of nine attack categories, following a standard IDS taxonomy.
The dataset is published in Hugging Face datasets format with explicit… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/cicflow-ids-multiclass.human_multi_classificationscicflow-ids-multiclass
CICFlow Multiclass Intrusion Detection Dataset
Overview
This dataset provides a multiclass network intrusion detection (IDS) benchmark derived from CICFlowMeter flow-level features.
It is designed for attack-type classification, robust IDS research, and interpretable security modeling.
Each network flow is labeled as either benign or one of nine attack categories, following a standard IDS taxonomy.
The dataset is published in Hugging Face datasets format with explicit… See the full description on the dataset page: https://huggingface.co/datasets/DollarSign/cicflow-ids-multiclass.multi_class_solidity_function_vulnerabilty
Dataset Card for "multi_class_solidity_function_vulnerabilty"
More Information needed
gpt_4o_mini_classifications_multi_humanqwen2_72b_classifications_multi_humanllama_3_1_8b_classifications_multi_humansafety-llama-multiclassbbc_news_multiclass_train_val_testLabel Names:
{
'business': 0,
'entertainment': 1,
'politics': 2,
'sport': 3,
'tech': 4
}
Dataset: Kaggle - BBC Full Text Document Classification
stack_exchange_multiclass_max_500diverseVul-multi-classossetian-news-multiclass
Dataset Card for Ossetian News Multiclass
Dataset Details
Dataset Description
The Ossetian News Multiclass dataset contains 13,036 manually curated Ossetian news articles categorized into 12 thematic classes. It is designed for multi‑class text classification of news in the Ossetian language, a low‑resource language spoken in the Caucasus. Each record includes the full news text, a single thematic label (provided in Russian), and the original source… See the full description on the dataset page: https://huggingface.co/datasets/OssetianNLPWorld/ossetian-news-multiclass.gpt_4o_classifications_multi_humanmistral_large_classifications_multi_humangemini_1_5_pro_classifications_multi_human
