datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-sentiment-classification
MultilingualSentimentClassification
An MTEB dataset
Massive Text Embedding Benchmark
Sentiment classification dataset with binary
(positive vs negative sentiment) labels. Includes 30 languages and dialects.
Task category
t2c
DomainsReviews, Written
Reference
https://huggingface.co/datasets/mteb/multilingual-sentiment-classification
How to evaluate on this task
You can evaluate an embedding model on this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-sentiment-classification.multilingual-scala-classification
ScalaClassification
An MTEB dataset
Massive Text Embedding Benchmark
ScaLa a linguistic acceptability dataset for the mainland Scandinavian languages automatically constructed from dependency annotations in Universal Dependencies Treebanks.
Published as part of 'ScandEval: A Benchmark for Scandinavian Natural Language Processing'
Task category
t2c
Domains
Fiction, News, Non-fiction, Blog, Spoken, Web, Written
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/multilingual-scala-classification.multilingual-safety-classification-dataset
Multilingual Safety Classification Dataset
A multilingual dataset for safety classification across 60 languages, created by Hasan Kurşun through machine translation of English safety prompts using NLLB-200-3.3B.
Dataset Details
Processed by: Hasan KurşunAuthor: Hasan KurşunYear: 2025Source Dataset: mvrcii/safety-moderation-benchmarkTranslation Model: facebook/nllb-200-3.3B
Languages (60)
African Languages (16): Amharic, Hausa, Kinyarwanda, Luganda… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/multilingual-safety-classification-dataset.multilingual-sentiment-vie-classification
MultiLingualSentiment_vie_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-vie-classification.
Splits
split
rows
test
684
train
2,358
validation
330
multilingual-sentiment-tha-classification
MultiLingualSentiment_tha_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-tha-classification.
Splits
split
rows
test
2,344
train
8,102
validation
1,153
toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/malexandersalazar/toxicity-multilingual-binary-classification-dataset.multilingual-sentiment-ind-classification
MultiLingualSentiment_ind_Classification
Deduplicated copy of kornwtp/multilingual-sentiment-ind-classification.
Splits
split
rows
test
2,266
train
7,926
validation
1,132
multilingual-classification-0001
Multilingual Text Classification Dataset
This dataset is designed for multilingual text classification tasks.
It includes labeled text samples across 8 languages, making it ideal for training and evaluating models on cross-lingual transfer, language identification, and multilingual understanding.
Dataset Overview
Split
# Examples
Size (bytes)
Train
18,657
2,651,248
Validation
2,665
378,709
Test
5,331
757,560
Total
26,653
3,787,517
Total… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/multilingual-classification-0001.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxicity-multilingual-binary-classification-dataset.multilingual-sentiment-ind-classificationMultilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.multilingual-sentiment-vie-classificationmultilingual-sentiment-tha-classificationmultilingual-document-classification
Multilingual Document Classification Dataset
This dataset contains 100,000 text passages across 100 non-English language-script pairs sourced from the agentlans/HuggingFaceFW-finetranslations-100-languages-sample collection.
Each original text passage is paired with its English translation and has been programmatically annotated with domain, writing genre, and educational classifications to facilitate cross-lingual classification and domain adaptation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-document-classification.multilingual-sentiment-vie-classificationmultilingual-sentiment-tha-classificationmultilingual-sentiment-ind-classificationmteb-human-multilingual-sentiment-classification
Multilingual Sentiment Human Subset
Unified multilingual dataset with lang column (eng, ara, nor, rus).
Splits:
test: All languages combined
eng/ara/nor/rus: Individual language subsets
Multilingual-Financial-Sentiment-and-Intent-Classificationmultilingual-research-classification
Multilingual Scientific Text Classification Dataset (MAG FoS L1)
Overview
This dataset contains multilingual scientific text samples (Catalan, Spanish, and English) extracted from scientific publications.Each sample is labeled using Microsoft Academic Graph (MAG) Field of Study — Level 1 categories.
For each publication, the text field is a random selection of:
the title
the abstract
the title followed by the abstract (title + ". " + abstract)
This introduces… See the full description on the dataset page: https://huggingface.co/datasets/nicolauduran45/multilingual-research-classification.multilingual_customer_support_intent_classificationmulti-lingual-classification-v2multi-lingual-classification-v1multi-lingual-classification-v0
