datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
veracier-industries
EDiTh — Enterprise Digital Twin Benchmark
What is this dataset?
EDiTh (Enterprise Digital Twin) is an open benchmark for evaluating
enterprise search and RAG systems on documents that actually look like
the ones you deal with every day: multilingual, scanned, cross-referenced,
and full of the edge cases that break demos.
At its core is Véracier Industries S.A., a fictional but rigorously
grounded €1.8 B French industrial group: 7 subsidiaries across 5 countries
(France… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/veracier-industries.gpqa_diamond_multilingual
GPQA Diamond Multilingual
gpqa_diamond_multilingual is a multilingual version of the benchmark GPQA Diamond, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a graduate-level multiple-choice question in biology, physics, or chemistry, written and validated by domain experts, translated into the five target languages.
This release is a corrected version of shanchen/gpqa_diamond_mc_multilingual that fixes translation artifacts and errors.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/gpqa_diamond_multilingual.mgsm-rev2
MGSM-Rev2
MGSM-Rev2 is a corrected version of the MGSM benchmark, which evaluates multilingual mathematical reasoning on grade-school word problems across 10 languages.
Please refer to the original repository for details.
aime24_multilingual
AIME24 Multilingual
aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages.
This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.aime25_multilingual
AIME25 Multilingual
aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages.
This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.
