datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multilingual-CoT-Collection"""
_LICENSE = "CC BY 4.0"
_HOMEPAGE = "https://github.com/kaistAI/CoT-Collection"
_LANGUAGES = {
"ko": "Korean",
"fr": "French",
"ru": "Russian",
"ja": "Japanese",
"zh": "Chinese",
}
# _ALL_LANGUAGES = "all_languages"
class CoTCollectionMultiConfig(datasets.BuilderConfig):RAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.ai-culture-multilingual-json-dolma
AI-Culture Multilingual JSON + DOLMA Corpus
16M words · 12 languages · CC-BY-4.0
The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality.
This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.aime24_multilingual
AIME24 Multilingual
aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages.
This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.aime25_multilingual
AIME25 Multilingual
aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages.
This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors.
It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.speculators-multilingual-en-fr-de-it-es
Speculators Multilingual SFT Dataset (en/fr/de/it/es)
A multilingual instruction-following dataset in ShareGPT format, built to train draft models for speculative decoding across English, French, German, Italian and Spanish.
Summary
An English instruction-tuning corpus with part of it kept in English and the rest machine-translated into French, German, Italian and Spanish using tencent/Hunyuan-MT-7B. Provided as a single mixed-language, ShareGPT-formatted dataset… See the full description on the dataset page: https://huggingface.co/datasets/Infomaniak-AI/speculators-multilingual-en-fr-de-it-es.
