datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vdr-multilingual-trainaime_2025_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2025_multilingual.aime26-multilingualAIME2025-Multilingual
Description
This repository contains a multi language version of the AIME2025 dataset.
As the english reference version, we haved used the one created by the authors of MathArena.
For completness, we have included the english version also in this repository, please, refer to the one contained in the MathArena github repository for the original one (https://github.com/eth-sri/matharena/tree/main/data/aime). Many thanks to Jasper Dekoninck for the help in understanding the structure… See the full description on the dataset page: https://huggingface.co/datasets/fedric95/AIME2025-Multilingual.vdr-multilingual-train-corpusaime25-multilingualRAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.LongBench-multilingualWIP, please don't use yet
ai-culture-multilingual-json-dolma
AI-Culture Multilingual JSON + DOLMA Corpus
16M words · 12 languages · CC-BY-4.0
The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality.
This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.tasklist-grok4-multilingual-50000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-org
Run Parameters
Parameter
Value
Model
grok-4-1-fast-non-reasoning
Temperature
0.75
Total Tasks
50000
Concurrency
8 workers
API Base
https://api.x.ai/v1
Generated
2026-04-07 09:04:57
Budget Cap
$15.0000
Multilingual
Yes (en, de, fr, es, nl, zh, ar, ru)
Language Distribution
Language
Code
Tasks
Arabic
ar
6111
Chinese
zh
6058
German
de
6057
Spanish
es
6020… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok4-multilingual-50000x-unfiltered.multilingual-code-comments-fixed-8
Fixed-8
Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged.
Removed ID
Replacement ID
8000_5
1056_0
8000_14
4357_8
8000_15
4848_9
8000_16
29069_13
8000_17
1385_4
8000_18
5142_0
All 500 Greek rows now have all five expert accuracy labels. Original… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.multilingual-CulturalBench-Hardtasklist-grok-multilingual-100000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-org
Run Parameters
Parameter
Value
Model
grok-4-1-fast-reasoning
Temperature
0.9
Total Tasks
83052
Concurrency
30 workers
API Base
https://api.x.ai/v1
Generated
2026-04-07 14:31:14
Budget Cap
$15.0000
Multilingual
Yes (en, de, fr, es, nl, zh, ar, ru)
Language Distribution
Language
Code
Tasks
Arabic
ar
10446
German
de
10397
Dutch
nl
10353
Spanish
es
10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.multilingual_examplesponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.aime_2024_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2024_multilingual.aime24_multilingual
AIME24 Multilingual
aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages.
This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.RULER-multilingualaime25_multilingual
AIME25 Multilingual
aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages.
This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.vdr-multilingual-train-hn-mineMATH-multilingual-traincardiffnlp_tweet_sentiment_multilingual_translatedaiw_easy_multilingualMultilingual-MATH-500multilingual_toxicity_datasetaiw_hard_multilingualmultilingual-dilemmamultilingual-code-comments
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
This dataset helps us understand how Large Language Models (LLMs) can create code comments in different languages. While LLMs are good at coding tasks in English, we don't know much about how well they work in other languages. This dataset, along with our research, studies how LLMs generate code comments in English, Chinese, Dutch, Polish, and Greek. In our case, we have… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments.open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.multilingual-code-comments-fixed-7
