CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /Multilingual-Thinking Dataset summary Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model. This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models. You can load the dataset using: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.texttext-generation1K<n<10K118 likes9.2k downloads1y agoHugging Face02baber /multilingual_mmluMMLU professionally translated into 14 languages using professional human translators, sourced from OpenAI's simple-eval. Original files: english: https://openaipublic.blob.core.windows.net/simple-evals/mmlu.csv multilingual: https://openaipublic.blob.core.windows.net/simple-evals/mmlu_{language}.csv where language one of "AR-XY", "BN-BD", "DE-DE", "ES-LA", "FR-FR", "HI-IN", "ID-ID", "IT-IT", "JA-JP", "KO-KR", "PT-BR", "ZH-CN", "SW-KE", "YO-NG", "EN-US" texttext-generation100K<n<1M1 likes7.6k downloads2y agoHugging Face03Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes3.9k downloads2y agoHugging Face04nvidia /Nemotron-SFT-Multilingual-v1 Dataset Description: Nemotron-Multilingual-v1 is a multilingual reasoning dataset made by translating a subsample of SFT data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1 into to 6 languages (German, French, Japanese, German, Italian, Japanese, Chinese).The original datasets were translated with Qwen2.5-14B-Instruct, then filtered with heuristics to remove translation failures and hallucinations. The STEM subsets are further post-edited with an… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v1.text-generation18 likes3.1k downloads7mo agoHugging Face05nvidia /Nemotron-SFT-Multilingual-v2 Dataset Description: Nemotron-SFT-Multilingual-v2 is a multilingual supervised fine-tuning (SFT) dataset for post-training text-generation models. It is generated by translating seed data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1, adding multilingual coverage for Hindi (hi), Korean (ko), Brazilian Portuguese (pt-br), and refreshed Japanese (ja) data. The dataset is generated with a new data processing pipeline that avoids line-breaking… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v2.texttext-generation100K<n<1M14 likes2.8k downloads4mo agoHugging Face06ellamind /gsm8k-platinum-multilingual GSM8K Platinum Multilingual Multilingual translations of GSM8K Platinum, a rigorously cleaned and verified version of GSM8K containing 1,209 elementary math word problems requiring multi-step arithmetic reasoning. Source: madrylab/gsm8k-platinum (test split, 1,209 questions) Languages Config Language Examples ces Czech 100 dan Danish 100 deu German 1,209 fin Finnish 100 fra French 100 ita Italian 100 nld Dutch 100 pol Polish 100 spa Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gsm8k-platinum-multilingual.textquestion-answering1K<n<10K1 likes1.7k downloads6mo agoHugging Face07ellamind /hle-multilingual HLE Multilingual Multilingual translations of HLE (Humanity's Last Exam), an expert-level QA benchmark with questions across math, science, humanities, and engineering designed to challenge even domain experts. Source: cais/hle (test split, 2,158 text-only questions out of 2,500 total) Languages Config Language Examples ces Czech 50 dan Danish 50 deu German 800 fin Finnish 50 fra French 50 ita Italian 50 nld Dutch 50 pol Polish 50 spa… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/hle-multilingual.textquestion-answering1K<n<10K0 likes1.5k downloads7mo agoHugging Face08giuliolovisotto /openai_multilingual_mmluMMLU professionally translated into 14 languages using professional human translators, sourced from OpenAI's simple-eval. Original files: english: https://openaipublic.blob.core.windows.net/simple-evals/mmlu.csv multilingual: https://openaipublic.blob.core.windows.net/simple-evals/mmlu_{language}.csv where language one of "AR-XY", "BN-BD", "DE-DE", "ES-LA", "FR-FR", "HI-IN", "ID-ID", "IT-IT", "JA-JP", "KO-KR", "PT-BR", "ZH-CN", "SW-KE", "YO-NG", "EN-US" texttext-generation100K<n<1M1 likes1.3k downloads2y agoHugging Face09lightonai /Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens. It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.texttext-generation1M<n<10M2 likes1.2k downloads4mo agoHugging Face10CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads12d agoHugging Face11aryashah00 /multilingual-sycophancy Multilingual Sycophancy A Parallel Benchmark for Cross-Lingual Alignment Failure across 38 Languages, 33 Opinion Categories, and 3 Resource Tiers. This dataset accompanies the research paper Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models. It contains 188,100 parallel records (4,950 per language × 38 languages) — each a triple of (prompt, sycophantic response, non-sycophantic response) — designed for forced-choice… See the full description on the dataset page: https://huggingface.co/datasets/aryashah00/multilingual-sycophancy.texttext-classification100K<n<1M0 likes738 downloads2mo agoHugging Face12eddie-OB /gsm8k-multilingual-reasoning gsm8k-multilingual-reasoning GSM8K with reasoning translated to multiple languages Schema {"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}} Usage from datasets importload_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K1 likes675 downloads8mo agoHugging Face13argilla /databricks-dolly-15k-curated-multilingual Dataset Card for "databricks-dolly-15k-curated-multilingual" A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below. STATUS: Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.texttext-generation10K<n<100K54 likes519 downloads3y agoHugging Face14nhagar /c4_urls_multilingual Dataset Card for c4_urls_multilingual This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (multilingual variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_multilingual.texttext-generation1B<n<10B1 likes467 downloads1y agoHugging Face15Similoluwa /african-multilingual-tokenizer-challenge African Multilingual Tokenizer Challenge dataset The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split. Split Per language Total Train 40,000 240,000 Validation 4,000 24,000 Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.texttext-generation100K<n<1M0 likes415 downloads21d agoHugging Face16eddie-OB /gsm8k-multilingual gsm8k-multilingual GSM8K translated to multiple languages (no reasoning) Schema {"prompt": "...", "answer": "...", "metadata": {...}} Usage from datasets import load_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K0 likes411 downloads8mo agoHugging Face17agentlans /multilingual-text Multilingual Text Dataset This dataset contains a curated selection of rows from multiple input datasets, where each row includes a text chunk of approximately 2000 tokens (as measured by Llama 3.1 tokenizer) verified to be written in the correct language. Only rows with properly classified language chunks are retained, ensuring high-quality multilingual data for analysis or model training. Preprocessing Steps Normalized whitespace, punctuation, Unicode characters, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-text.texttext-generation1M<n<10M5 likes401 downloads1y agoHugging Face18ellamind /gpqa-multilingualgated GPQA Multilingual Multilingual translations of GPQA (Graduate-Level Google-Proof Q&A), a challenging multiple-choice benchmark requiring graduate-level expertise in biology, physics, and chemistry. Source: Idavidrein/gpqa (gpqa_main, 448 questions) Languages Config Language Examples ces Czech 448 dan Danish 448 deu German 448 fin Finnish 50 fra French 448 ita Italian 448 nld Dutch 448 pol Polish 448 spa Spanish 448 More to be added later.… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gpqa-multilingual.textquestion-answering1K<n<10K0 likes389 downloads7mo agoHugging Face19agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes366 downloads2y agoHugging Face20dgallitelli /multilingual-wealth-alpaca Multilingual Wealth Alpaca Dataset Work derivative of gbharti/wealth-alpaca_lora dataset. The original dataset is a combination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5 . This version is a cleaned up version, which also has: mutlilingual support (en, it, fr, es, de) CSV and JSON files text-generation0 likes347 downloads3y agoHugging Face21kaist-ai /Multilingual-CoT-Collection""" _LICENSE = "CC BY 4.0" _HOMEPAGE = "https://github.com/kaistAI/CoT-Collection" _LANGUAGES = { "ko": "Korean", "fr": "French", "ru": "Russian", "ja": "Japanese", "zh": "Chinese", } # _ALL_LANGUAGES = "all_languages" class CoTCollectionMultiConfig(datasets.BuilderConfig):text-generation100K<n<1M28 likes346 downloads3y agoHugging Face22ellamind /simpleqa-verified-multilingual SimpleQA Verified Multilingual Multilingual translations of SimpleQA Verified, a 1,000-prompt factuality benchmark from Google DeepMind that evaluates short-form parametric knowledge (facts stored in model weights). Source: google/simpleqa-verified (eval split, 1,000 examples) Languages Config Language Examples ces Czech 100 dan Danish 100 deu German 1,000 fra French 100 ita Italian 100 nld Dutch 100 pol Polish 100 spa Spanish 100 More to… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual.textquestion-answering1K<n<10K1 likes343 downloads7mo agoHugging Face23Gabrui /multilingual_TinyStories Dataset Card for Multilingual TinyStories Dataset Details Dataset Description The Multilingual TinyStories dataset contains translations of the original TinyStories dataset, which consists of synthetically generated short stories using a small vocabulary suitable for 3 to 4-year-olds. These stories were originally generated by GPT-3.5 and GPT-4. The multilingual versions have been translated into various languages, including Spanish, Chinese, German, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/multilingual_TinyStories.texttext-generation10M<n<100M1 likes322 downloads2y agoHugging Face24risaleinur /risale-nur-multilingual Risale-i Nur Multilingual Corpus Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir. Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.tabulartranslation100K<n<1M2 likes313 downloads28d agoHugging Face25Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes282 downloads2y agoHugging Face26Dxniz /TinyStories-Multilingual Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes. The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.texttext-generation10K<n<100K1 likes281 downloads6mo agoHugging Face27deokhk /multilingual_reasoning_gap_outputs Dataset Card for multilingual_reasoning_gap_outputs Paper | Code Dataset Details Dataset Description This dataset contains experiment outputs for Qwen3-4B used in our study on multilingual reasoning gaps. It includes: Prober checkpoints trained for understanding-failure analysis Intermediate results, such as: Model inference outputs Signals for understanding failure detection Auxiliary artifacts used for probing and analysis The dataset is released to… See the full description on the dataset page: https://huggingface.co/datasets/deokhk/multilingual_reasoning_gap_outputs.text-generation0 likes278 downloads9mo agoHugging Face28agentlans /multilingual-sentences Multilingual Sentences Dataset contains sentences from 50 languages, grouped by their two-letter ISO 639-1 codes. The "all" configuration includes sentences from all languages. Dataset Overview Multilingual Sentence Dataset is a comprehensive collection of high-quality, linguistically diverse sentences. Dataset is designed to support a wide range of natural language processing tasks, including but not limited to language modeling, machine translation, and cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-sentences.texttext-generation10M<n<100M6 likes268 downloads2y agoHugging Face29yagizgencer /emojinize-multilingual Emojinize Multilingual A multilingual dataset of 108,478 sentences across 14 languages for emoji-based text augmentation. In each sentence, selected spans (individual words or fixed multi-word expressions) are identified by character offsets and paired with emoji sequences representing their meaning in context. The dataset supports downstream span detection and emoji generation tasks, and was created using a two-stage LLM annotation pipeline with gpt-5.4 for span marking and… See the full description on the dataset page: https://huggingface.co/datasets/yagizgencer/emojinize-multilingual.texttoken-classification100K<n<1M2 likes232 downloads2mo agoHugging Face30textdetox /multilingual_paradetoxMultilingual Text Detoxification with Parallel Data This is the multilingual parallel dataset for the text detoxification task. Prepared for TextDetox Shared Task. 📰 Updates [2025] The second edition of TextDetox shared task! webpage [2025] We extend our data to new languages! Now also included: Italian, French, Hebrew, Hinglish, Japanese, Tatar. Check our test part. [2025]We dived into the explainability of our data in our new COLING paper! [2024] You can check additional releases for… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_paradetox.texttext-generation1K<n<10K11 likes222 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.