datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiPICo
Dataset Summary
MultiPICo (Multilingual Perspectivist Irony Corpus) is a disaggregated multilingual corpus for irony detection, containing 18,778 pairs of short conversations (post-reply) from Twitter (8,956) and Reddit (9,822), along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/MultiPICo.multilingual-hatespeech-dataset
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset
Description
This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate
texts also the data from different languages needed to be identified as a corresponding
correct language. The following are the languages in the dataset with the numbers corresponding to that language.
(1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.EPIC
Dataset Card for EPICorpus
Dataset Summary
EPIC (English Perspectivist Irony Corpus) is a disaggregated English corpus for irony detection, containing 3,000 pairs of short conversations (posts-replies) from Twitter and Reddit, along with the demographic information of each annotator (age, nationality, gender, and so on).
Supported Tasks and Leaderboards
Irony classification task using soft labels (i.e., distribution of annotations) or hard labels… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Perspectivist-NLU/EPIC.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.multilingual_evalsmultilingual-frequency-lists
Multilingual Frequency Lists
This dataset contains multiple word-frequency lists in various languages such as French, Japanese, Spanish, Italian and Portuguese.
Specifically, these are frequency lists of lemmas, meaning, for example, that words like 'run', 'runs' and 'running' are counted together as occurences of the same lemma 'run'.
These frequency lists were generated from ~1GB of subtitles scraped from a variety of Netflix shows and films and parsed using relevant spacy models… See the full description on the dataset page: https://huggingface.co/datasets/joshdavham/multilingual-frequency-lists.multilingual-islr-mediapipe
Multilingual ISLR MediaPipe Landmarks
Dataset Description
This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning.
The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.korea-places-multilingual
Korean Place Names, Multilingual
Built and maintained by Korea Basics, a sourced guide to Korean
entry rules and getting around, published in seven languages.
16,126 places in South Korea with their Korean (Hangul) name next to the
romanized English name, plus Japanese and Chinese names where the source has
them, coordinates, road-name address, and subway lines for stations.
Why this exists
A visitor who reads "Gyeongbokgung Palace" in a guide cannot type that… See the full description on the dataset page: https://huggingface.co/datasets/hjm1980/korea-places-multilingual.phishing-emails-multilingual
Phishing Emails Multilingual (ID/EN) — Synthetic
Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah.
⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal.
Ringkasan
600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID
Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es
Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs
The code for processing can be found here
Useful for data distillation, training or benchmarking.
Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.Nemotron-SFT-Multilingual-v2-prompt-only
Nemotron-SFT-Multilingual-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v2-prompt-only.quran_multilingual_parallel
📘 Qur’an Multilingual Parallel Dataset (quran_multilingual_parallel)
This dataset presents a clean, structurally-aligned multilingual parallel corpus of the Qur’anic text. It is intended for linguistic, computational, and cross-lingual AI applications — not only for religious interpretation.
It contains over 6,200 verse-level alignments in 54 human languages, formatted in a machine-friendly .csv structure with language-specific translation fields.
🧠 Dataset Highlights… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran_multilingual_parallel.sarcasm_headlines_multilingual
Dataset Card for Multilingual Sarcasm Detection
Dataset Summary
Dataset consists of news article headlines in Dutch, English and Italian. The news article headlines are both from actual news sources and sarcastic/satirical newspapers. The news article is determined sarcastic/non-sarcastic based on the news article source.
The sources of news articles are:
The Huffington Post (en, non-sarcastic)
The Onion (en, sarcastic)
NOS (nl, non-sarcastic)
De Speld (nl, sarcastic)
Il… See the full description on the dataset page: https://huggingface.co/datasets/helinivan/sarcasm_headlines_multilingual.Nemotron-SFT-Multilingual-v1-prompt-only
Nemotron-SFT-Multilingual-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v1-prompt-only.luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.gibberish_multilingualred_team_agent_analysis_multilingual_story_analysis_detailed
red_team_agent_analysis_multilingual_story_analysis_detailed
This dataset was automatically uploaded from the red-team-agent repository.
Dataset Information
Original file: multilingual_story_analysis_detailed.csv
Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis/multilingual_story_analysis_detailed.csv
Validation: Valid CSV with 1000 rows, 12 columns (0.1MB)
Usage
import pandas as pd
from datasets import load_dataset
# Load using datasets… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_multilingual_story_analysis_detailed.multilingual-EMIR-reporting-csvmultilingual_abusive-non-abusivemultilingualcrowspairs
[!NOTE]
Dataset origin: https://gitlab.inria.fr/corpus4ethics/multilingualcrowspairs/
MultiLingualCrowsPairs
Multilingual CrowS-Pairs, a challenge dataset for measuring stereotypical biases present in the masked language models (MLMs) in 7 different languages.
This challenge dataset was built on the Crows-Pairs corpus (Nangia et al. 2020) using the methodology described in (Névéol et al. 2023).
The 7 new languages are the following:
Arabic from Maghreb and the Arab world in… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingualcrowspairs.multilingual_jailbreak_challengesProcessed_TTS_Multilingual_Data
Processed TTS Multilingual Data
Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages.
Datasets Included
Subset
Samples
Hours
Description
indic_voices_r
239,684
548.8h
Indic Voices_R — IVR recordings
rasa
201,509
361.2h
RASA — read speech (wiki, conv, book, news)
indictts_iitm
155,236
253.6h
Indic TTS (IIT Madras) — studio TTS recordings at 48kHz
Total
596,429
1,163.6h
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.do_not_answer_multilingual_12scienceqa-multilingual-hindi#ScienceQA Hindi Translation Dataset
##Dataset Description
This dataset is a Hindi-translated version of the original ScienceQA dataset. It includes multiple-choice science questions, with fields for:
Images (optional visual context),
Hints (optional support text),
English questions and their Hindi translations,
Multiple answer choices,
Correct answers.
This translation is intended to support multilingual education research, question-answering in Hindi, and fairness studies in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/model2me/scienceqa-multilingual-hindi.MultilingualTranscriptionDataset
MultilingualTranscriptionDataset
tags: Transcription, LanguageProcessing, Multilingual
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MultilingualTranscriptionDataset' is a curated collection of text transcriptions from various audio recordings. Each transcription is provided in multiple languages, emphasizing the diversity and complexity of language processing. This dataset aims to assist in developing machine learning… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MultilingualTranscriptionDataset.multi-lingual-llm
Dataset Card for Dataset Name
This data set contains set questions in tamil and possible answers, with the correct answer in the column. This helps to test LLM to see for accuracy.
Dataset Details
Dataset Description
Curated by: Madhumitha Sivalingapandian
Language(s) (NLP): Tamil and English
License: [More Information Needed]
Uses
Used for measuring performance of LLM.
Direct Use
Spoken language accuracy measurement dataset… See the full description on the dataset page: https://huggingface.co/datasets/MithuSi/multi-lingual-llm.opus100-multilingualMultilingual_E-commerce
