datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
task1577_amazon_reviews_multi_japanese_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1577_amazon_reviews_multi_japanese_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1577_amazon_reviews_multi_japanese_language_classification.Multi-Lingual-Lyrics-for-Genre-Classificationfrom https://www.kaggle.com/datasets/mateibejan/multilingual-lyrics-for-genre-classification
task638_multi_woz_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task638_multi_woz_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task638_multi_woz_classification.synthetic-text-classification-news-multi-label
Dataset Card for synthetic-text-classification-news-multi-label
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/synthetic-text-classification-news-multi-label/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-text-classification-news-multi-label.multi_edu_classificationhuman_multi_classifications_500human_multi_classificationsinappropriateness-token-classification-binarized-multi-refamazon_reviews_multi_fr_prompt_stars_classification
amazon_reviews_multi_fr_prompt_stars_classification
Summary
amazon_reviews_multi_fr_prompt_stars_classification is a subset of the Dataset of French Prompts (DFP).It contains 4,620,000 rows that can be used for a stars-classification sentiment analysis task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_stars_classification.task1575_amazon_reviews_multi_sentiment_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1575_amazon_reviews_multi_sentiment_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1575_amazon_reviews_multi_sentiment_classification.amazon_reviews_multi_fr_prompt_classes_classification
amazon_reviews_multi_fr_prompt_classes_classification
Summary
amazon_reviews_multi_fr_prompt_classes_classification is a subset of the Dataset of French Prompts (DFP).It contains 4,480,000 rows that can be used for a text classification task.The original data (without prompts) comes from the dataset amazon_reviews_multi by Keung et al. where only the French split has been kept.
A list of prompts (see below) was then applied in order to build the input and target columns… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/amazon_reviews_multi_fr_prompt_classes_classification.gpt_4o_mini_classifications_multi_humanllama_3_1_8b_classifications_multi_humanqwen2_72b_classifications_multi_humanmulti_class_classification_datasetinappropriateness-token-classification-multi-refgpt_4o_classifications_multi_humanmistral_large_classifications_multi_humantask1576_amazon_reviews_multi_english_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1576_amazon_reviews_multi_english_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1576_amazon_reviews_multi_english_language_classification.llama_3_1_405b_classifications_multi_humangemini_1_5_pro_classifications_multi_humanflan_combined_task1576_amazon_reviews_multi_english_language_classificationllama_3_1_70b_classifications_multi_humancommand_r_plus_classifications_multi_humanmulti-lingual-classification-v2flan_combined_task1575_amazon_reviews_multi_sentiment_classificationmars-multi-label-classification
MER - Mars Exploration Rover Dataset
A multi-label classification dataset containing Mars images from the Mars Exploration Rover (MER) mission for planetary science research.
Dataset Metadata
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Version: 1.0
Date Published: 2025-05-10
Cite As: TBD
Classes
This dataset uses multi-label classification, meaning each image can have multiple class labels.
The dataset contains the following classes:… See the full description on the dataset page: https://huggingface.co/datasets/gremlin97/mars-multi-label-classification.multi_class_classification_dataset_openai_predictionsclassification-medicale-multi-cancer
