datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dbarbedillo/SMS_Spam_Multilingual_Collection_Dataset.conceptual-12m-mbart-50-multilingualSMS_Spam_Multilingual_Collection_DatasetSMS Spam Multilingual Collection Dataset
Collection of Multilingual SMS messages tagged as spam or legitimate
About Dataset
Context
The SMS Spam Collection is a set of SMS-tagged messages that have been collected for SMS Spam research. It originally contained one set of SMS messages in English of 5,574 messages, tagged according to being ham (legitimate) or spam and later Machine Translated into Hindi, German and French.
The text has been further translated into Spanish, Chinese, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/KumarSahil299885/SMS_Spam_Multilingual_Collection_Dataset.conceptual-12m-multilingual-marianThis dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models. Data distribution is following:
train_file_marian_final.tsv: 10010625 captions (2502656 captions of English, German, Spanish, French each)
val_file_marian_final.tsv:… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian.Indian-Multilingual-Bias-Dataset
Indian Multilingual Bias Dataset
Dataset Description
The Indian Multilingual Bias Dataset is a comprehensive collection designed to evaluate and measure social biases in Large Language Models (LLMs) across three major Indian languages: English, Bengali (বাংলা), and Hindi (हिंदी). This dataset is based on the original Indian-BhED dataset and focuses on four critical dimensions of bias prevalent in Indian society.
Key Features
🌐 Multilingual:… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Indian-Multilingual-Bias-Dataset.Multilingual-Needle-in-a-Haystack
Multilingual Needle in a Haystack (MLNeedle)
The MultiLingual Needle-in-a-Haystack (MLNeedle) test is a dataset designed to assess how well Large Language Models (LLMs) find specific information ("needle") within long, multilingual texts ("haystack"). Built on MLQA, it contains over 5,000 extractive question-answer instances across seven languages (English, Arabic, German, Spanish, Hindi, Vietnamese, Simplified Chinese). We systematically vary the "needle's" language and position to… See the full description on the dataset page: https://huggingface.co/datasets/ameyhengle/Multilingual-Needle-in-a-Haystack.MultiPICo
Dataset Summary
MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.conceptual-12m-multilingual-marian-128This dataset is created from subset of Conceptual Captions. The original dataset has 12M captions but this dataset has around 10M image, caption pairs in different languages with 2.5M unique images. This dataset has captions translated from English to Spanish, German, French using language specific English to Marian models (with sequence length 128). Data distribution is following:
train_file_marian_final.tsv: 10002432 captions (2500608 captions of English, German, Spanish, French each)… See the full description on the dataset page: https://huggingface.co/datasets/flax-community/conceptual-12m-multilingual-marian-128.multilingual-hatespeech-dataset
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset
Description
This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate
texts also the data from different languages needed to be identified as a corresponding
correct language. The following are the languages in the dataset with the numbers corresponding to that language.
(1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.EPIC
Dataset Card for EPICorpus
Dataset Summary
EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.multilingual-lima
Multilingual LIMA
A multilingual extension of the LIMA instruction-tuning dataset. The original English prompt–response pairs were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
User instruction (translated; en is the original).
output
Assistant response (translated; en is the original).
Languages (configs): en (original), zh, it, bn… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-lima.multilingual-elder-safety-msgs
multilingual-elder-safety-msgs
A hand-authored, multilingual elder fraud-recognition and safety coaching dataset. 467 curated scam/safe scenarios in Chinese and English, with platform-generated coaching responses localized across 5 languages: Chinese, English, Vietnamese, Khmer (Cambodian), and Lao. Expanded to 1,029 rows through Adaption Labs platform reasoning traces and multilingual adaptation.
Built for communities where filial piety, authority deference, and fear of… See the full description on the dataset page: https://huggingface.co/datasets/vanila434/multilingual-elder-safety-msgs.multilingual-s1
Multilingual s1
A multilingual extension of the s1K-1.1 reasoning dataset. The original English reasoning questions and DeepSeek-R1 distilled solutions were translated into 9 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
We filter the upstream simplescaling/s1K-1.1 corpus to keep only samples whose DeepSeek-R1 trajectories were marked as correctly distilled, then translate the resulting subset.… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-s1.multilingual-vqasynthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.conceptual-12m-multilingual-marian-esDeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.afrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.multilingual-safety
Multilingual Safety Instructions
A multilingual extension of the safety-only instruction–refusal pairs released with the Safety-Tuned LLaMAs project. The original 1,000 harmful-prompt / refusal-response pairs (English) were translated into 11 additional typologically diverse languages with google/gemini-2.0-flash-001. Each language is stored as a separate Hugging Face config.
Field
Description
prompt
Harmful user instruction (translated; en is the original).
output
Safe… See the full description on the dataset page: https://huggingface.co/datasets/iNLP-Lab/multilingual-safety.multilingual_evalsKinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English.
The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples.
multilingual-financial-sentiment
Multilingual Financial Sentiment Dataset
A curated dataset of 39,829 financial news sentences annotated with sentiment labels (Negative / Neutral / Positive) across 7 languages, collected from 80+ financial news sources worldwide.
Dataset Summary
Total samples
39,829
Languages
7 (EN, ZH, JA, DE, FR, ES, AR)
Labels
3 (negative, neutral, positive)
Format
CSV
Sources
80+ financial news outlets
Languages
Language
Code
Samples
%… See the full description on the dataset page: https://huggingface.co/datasets/Kenpache/multilingual-financial-sentiment.Multilingual_medical_symptom_triage
tags:
- medical
- healthcare
- classification
- outbreak-detection
- triage
- multilingual
- adaption
- india
Multilingual Medical Symptom Triage Dataset
Dataset Description
A Mutlilingual medical triage dataset containing 9,064 patient
symptom descriptions in Hindi, English, and Hinglish (code-mixed
Hindi-English), paired with triage recommendations and rich
clinical metadata. Designed for training multilingual triage
classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.multilingual-frequency-lists
Multilingual Frequency Lists
This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese.
Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'.
These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.wikipedia-2023-11-kikongo-lingala-cohere-multilingual-v3multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.instructional-dialogues-multilingual
Multilingual Instructional Dialogues (10-Language Dataset)
Multilingual Instructional Dialogues is a high-quality dataset of 100 structured, goal-oriented dialogues in 10 major world languages, created for training and fine-tuning AI assistants, chatbots, and instruction-tuned large language models.
Each dialogue simulates a clear, polite interaction where a user asks for guidance on how to perform a task, and the assistant responds with easy-to-follow steps. This dataset has been… See the full description on the dataset page: https://huggingface.co/datasets/Raftico/instructional-dialogues-multilingual.YDX07_Multilingual_Corpus_2026ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.Multilingual_Topic-Specific_Article-Extraction_and_Classification
Dataset Card for Multilingual Historical News Article Extraction and Classification Dataset
This dataset was created specifically to test Large Language Models' (LLMs) capabilities in processing and extracting topic-specific content from historical newspapers based on OCR'd text.
Cite the Dataset
Mauermann, Johanna, González-Gallardo, Carlos-Emiliano, and Oberbichler, Sarah. (2025). Multilingual Topic-Specific Article-Extraction and Classification [Data set]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/oberbics/Multilingual_Topic-Specific_Article-Extraction_and_Classification.
