CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaysia-ai /Multilingual-TTS Multilingual-TTS A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips. Each config is one source dataset normalized to rows of {audio_filename, text, speaker}: audio_filename — clip path inside that config's <config>_audio.zip (mono MP3). text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.24 likes1.1k downloads1mo agoHugging Face02malaysia-ai /Multilingual-TTS-language Multilingual-TTS-language malaysia-ai/Multilingual-TTS with two extra columns: column description audio_filename, text, speaker unchanged from malaysia-ai/Multilingual-TTS language language detected from the text (transcript) column of every row post-normalized text after rule-based punctuation / capitalization normalization All original columns and the file/folder layout are preserved: 1552 subsets / 1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.text-to-speech100M<n<1B0 likes785 downloads1mo agoHugging Face03nomic-ai /vdr-multilingual-trainimage100K<n<1M0 likes466 downloads2y agoHugging Face04shanchen /aime_2025_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy https://arxiv.org/abs/2505.22888 Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2025_multilingual.tabularn<1K0 likes371 downloads1y agoHugging Face05ellamind /aime26-multilingualtextn<1K0 likes359 downloads7mo agoHugging Face06fedric95 /AIME2025-Multilingual Description This repository contains a multi language version of the AIME2025 dataset. As the english reference version, we haved used the one created by the authors of MathArena. For completness, we have included the english version also in this repository, please, refer to the one contained in the MathArena github repository for the original one (https://github.com/eth-sri/matharena/tree/main/data/aime). Many thanks to Jasper Dekoninck for the help in understanding the structure… See the full description on the dataset page: https://huggingface.co/datasets/fedric95/AIME2025-Multilingual.tabularn<1K3 likes340 downloads10mo agoHugging Face07nomic-ai /vdr-multilingual-train-corpusimage100K<n<1M0 likes333 downloads2y agoHugging Face08kaist-ai /Multilingual-CoT-Collection""" _LICENSE = "CC BY 4.0" _HOMEPAGE = "https://github.com/kaistAI/CoT-Collection" _LANGUAGES = { "ko": "Korean", "fr": "French", "ru": "Russian", "ja": "Japanese", "zh": "Chinese", } # _ALL_LANGUAGES = "all_languages" class CoTCollectionMultiConfig(datasets.BuilderConfig):text-generation100K<n<1M28 likes321 downloads3y agoHugging Face09malaysia-ai /multipacking-multilingual-tts-10k-qwen3 multipacking-multilingual-tts-10k-qwen3 This is Mosaic format for https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS including multipacking max 10240 context length using Qwen3 tokenizer, so you can plug and play to train it. For training example, you can check https://huggingface.co/malaysia-ai/Qwen3-1.7B-Multilingual-TTS 0 likes261 downloads1y agoHugging Face10ellamind /aime25-multilingualtextn<1K0 likes228 downloads7mo agoHugging Face11projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes176 downloads2y agoHugging Face12AIML-TUDA /LongBench-multilingualWIP, please don't use yet text10K<n<100K1 likes120 downloads7mo agoHugging Face13AI-Culture-Commons /ai-culture-multilingual-json-dolma AI-Culture Multilingual JSON + DOLMA Corpus 16M words · 12 languages · CC-BY-4.0 The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality. This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.texttranslation1K<n<10K3 likes92 downloads1y agoHugging Face14empero-ai /tasklist-grok4-multilingual-50000x-unfiltered TaskGen Dataset Generated with taskgen by empero-org Run Parameters Parameter Value Model grok-4-1-fast-non-reasoning Temperature 0.75 Total Tasks 50000 Concurrency 8 workers API Base https://api.x.ai/v1 Generated 2026-04-07 09:04:57 Budget Cap $15.0000 Multilingual Yes (en, de, fr, es, nl, zh, ar, ru) Language Distribution Language Code Tasks Arabic ar 6111 Chinese zh 6058 German de 6057 Spanish es 6020… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok4-multilingual-50000x-unfiltered.tabular10K<n<100K2 likes81 downloads6mo agoHugging Face15AISE-TUDelft /multilingual-code-comments-fixed-8 Fixed-8 Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged. Removed ID Replacement ID 8000_5 1056_0 8000_14 4357_8 8000_15 4848_9 8000_16 29069_13 8000_17 1385_4 8000_18 5142_0 All 500 Greek rows now have all five expert accuracy labels. Original… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.text1K<n<10K0 likes74 downloads7d agoHugging Face16appier-ai-research /multilingual-CulturalBench-Hardtabular10K<n<100K0 likes72 downloads2y agoHugging Face17empero-ai /tasklist-grok-multilingual-100000x-unfiltered TaskGen Dataset Generated with taskgen by empero-org Run Parameters Parameter Value Model grok-4-1-fast-reasoning Temperature 0.9 Total Tasks 83052 Concurrency 30 workers API Base https://api.x.ai/v1 Generated 2026-04-07 14:31:14 Budget Cap $15.0000 Multilingual Yes (en, de, fr, es, nl, zh, ar, ru) Language Distribution Language Code Tasks Arabic ar 10446 German de 10397 Dutch nl 10353 Spanish es 10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.tabular100K<n<1M2 likes70 downloads6mo agoHugging Face18moonshine-ai /multilingual_examplesaudion<1K0 likes68 downloads1y agoHugging Face19wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes65 downloads25d agoHugging Face20wujoe132 /ponys-ai-multilingual-companion-evaluation Ponys.ai Multilingual AI Companion Evaluation Protocols This public collection contains 20 reusable evaluation protocols for AI companion and character experiences. It covers conversation memory, persona consistency, consent recovery, visual continuity, code switching, and regional language behavior across Japanese, Korean, Latin American Spanish, Brazilian Portuguese, Simplified Chinese, Traditional Chinese, and English. Each protocol includes structured metadata and a CSV… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-ai-multilingual-companion-evaluation.0 likes59 downloads2mo agoHugging Face21shanchen /aime_2024_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy https://arxiv.org/abs/2505.22888 Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2024_multilingual.textn<1K0 likes54 downloads1y agoHugging Face22lightonai /aime24_multilingual AIME24 Multilingual aime24_multilingual is a multilingual version of the benchmark AIME 2024, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2024, translated into the five target languages. This release is a corrected version of shanchen/aime_2024_multilingual that fixes translation artifacts and errors. It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime24_multilingual.textquestion-answeringn<1K0 likes53 downloads4mo agoHugging Face23AIML-TUDA /RULER-multilingualtabular10K<n<100K1 likes51 downloads6mo agoHugging Face24lightonai /aime25_multilingual AIME25 Multilingual aime25_multilingual is a multilingual version of the benchmark AIME 2025, covering six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample is a competition-level mathematics problem from the American Invitational Mathematics Examination (AIME) 2025, translated into the five target languages. This release is a corrected version of shanchen/aime_2025_multilingual that fixes translation artifacts and errors. It is released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/aime25_multilingual.textquestion-answeringn<1K0 likes51 downloads4mo agoHugging Face25nomic-ai /vdr-multilingual-train-hn-minetext100K<n<1M0 likes49 downloads2y agoHugging Face26appier-ai-research /MATH-multilingual-traintext1K<n<10K0 likes48 downloads2y agoHugging Face27AIM-Harvard /cardiffnlp_tweet_sentiment_multilingual_translatedtext1K<n<10K0 likes45 downloads2y agoHugging Face28shanchen /aiw_easy_multilingualtext1K<n<10K0 likes42 downloads2y agoHugging Face29appier-ai-research /Multilingual-MATH-500text1K<n<10K1 likes39 downloads2y agoHugging Face30AIM-Harvard /multilingual_toxicity_datasettext10K<n<100K0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.