datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench_Multilingual
SWE-bench Multilingual
Dataset Summary
SWE-bench Multilingual is a dataset that tests systems' ability to resolve real-world GitHub issues across a broad range of programming languages. The original SWE-bench is Python-only; this dataset extends the same task format to 9 languages drawn from 41 popular repositories.
The dataset collects 300 test Issue-Pull Request pairs. Evaluation is performed by unit test verification, using post-PR behavior as the reference solution.
The… See the full description on the dataset page: https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual.multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.Multilingual-Thinking
Dataset summary
Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model.
This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models.
You can load the dataset using:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.multilingual-NLI-26lang-2mil7
Datasheet for the dataset: multilingual-NLI-26lang-2mil7
Dataset Summary
This dataset contains 2 730 000 NLI text pairs in 26 languages spoken by more than 4 billion people. The dataset can be used to train models for multilingual NLI (Natural Language Inference) or zero-shot classification. The dataset is based on the English datasets MultiNLI, Fever-NLI, ANLI, LingNLI and WANLI and was created using the latest open-source machine translation models.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/MoritzLaurer/multilingual-NLI-26lang-2mil7.multilingual_librispeech_fr_processed
multilingual_librispeech_fr_processed
Dataset Description
Dataset Summary
The data files can be found on the illuin gcloud instance at this adress: unknown_url
This dataset has been processed from Huggingface Hub dataset facebook/multilingual_librispeech and the config french
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual_librispeech_fr_processed.multilingual
Dataset Card for "multilingual"
More Information needed
multilingual_cc_news
hotchpotch/multilingual_cc_news
Dataset Summary
This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language.
Source and transformation
Original source datasets on the Hugging Face Hub:
CloverSearch/cc-news-mutlilingual
intfloat/multilingual_cc_news
The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/multilingual_cc_news.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.wikipedia-2023-11-embed-multilingual-v3-int8-binary
Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings)
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.openalex-multilingual-embeddings
OpenAlex Multilingual Embeddings
This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20.
The dataset was created for the FORAS project to investigate the efficacy of
different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository.
The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.synthetic_pii_finance_multilingual
Image generated by DALL-E. See prompt for more details
💼 📊 Synthetic Financial Domain Documents with PII Labels
gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0.
This dataset is designed to assist with the following use cases:
🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.hle-multilingual
HLE Multilingual
Multilingual translations of HLE (Humanity's Last Exam), an expert-level QA benchmark with questions across math, science, humanities, and engineering designed to challenge even domain experts.
Source: cais/hle (test split, 2,158 text-only questions out of 2,500 total)
Languages
Config
Language
Examples
ces
Czech
50
dan
Danish
50
deu
German
800
fin
Finnish
50
fra
French
50
ita
Italian
50
nld
Dutch
50
pol
Polish
50
spa… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/hle-multilingual.SWE-Bench-MultilingualC_CPPFileteredgsm8k-platinum-multilingual
GSM8K Platinum Multilingual
Multilingual translations of GSM8K Platinum, a rigorously cleaned and verified version of GSM8K containing 1,209 elementary math word problems requiring multi-step arithmetic reasoning.
Source: madrylab/gsm8k-platinum (test split, 1,209 questions)
Languages
Config
Language
Examples
ces
Czech
100
dan
Danish
100
deu
German
1,209
fin
Finnish
100
fra
French
100
ita
Italian
100
nld
Dutch
100
pol
Polish
100
spa
Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gsm8k-platinum-multilingual.arc-multilingualSWE-Bench-MultilingualC_CPPFiletered_newmultilingual-TEDX-frThe french subset of the dataset Multilingual TEDx. The data uploaded to HF corresponds to the directory fr-fr. The audio files are automatically resampled to 16 kHz.
Configs:
single_samples (default): all samples taken separately
Sample
{'file': '0u7tTptBo9I-0', 'audio': {'path': None, 'array': array([ 3.05175781e-05, 6.10351562e-05, 9.15527344e-05, ...,
-2.44140625e-04, -3.35693359e-04, -2.74658203e-04]), 'sampling_rate': 16000}, 'sentence': "Bonsoir ! Notre… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual-TEDX-fr.Dolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens.
It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.tweet_topic_multilingual
Dataset Card for "cardiffnlp/tweet_topic_multilingual"
Dataset Summary
This is the official repository of X-Topic (Multilingual Topic Classification in X: Dataset and Analysis, EMNLP 2024), a topic classification dataset based on X (formerly Twitter), featuring 19 topic labels.
The classification task is multi-label, with tweets available in four languages: English, Japanese, Spanish, and Greek.
The dataset comprises 4,000 tweets (1,000 per language), collected between… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_topic_multilingual.humaneval-multilingualSWE-bench_Multilingualhellaswag-multilingualrag_multilingual_training_negatives
How this dataset was made
We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM.
We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document.
We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM.
Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.multilingual_toxicity_dataset
Multilingual Toxicity Detection Dataset
[2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task.
[2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language.
Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.
The… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_toxicity_dataset.Mercury_Multilingualnanobeir-multilingual-extended
NanoBEIR Multilingual Extended Dataset
This dataset extends the NanoBEIR multilingual collection with Japanese and Korean translations.
Dataset Structure
Each configuration follows the pattern <BASE>_<LANG> with splits:
corpus: Document corpus
queries: Search queries
qrels: Query relevance judgments (when available)
Languages
Arabic (ar), German (de), English (en), Spanish (es), French (fr)
Italian (it), Norwegian (no), Portuguese (pt), Swedish (sv)… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/nanobeir-multilingual-extended.
