multilingual-ai
Multilingual-TTS
Multilingual-TTS
A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips.
Each config is one source dataset normalized to rows of {audio_filename, text, speaker}:
audio_filename — clip path inside that config's <config>_audio.zip (mono MP3).
text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.vdr-multilingual-trainaime_2025_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/aime_2025_multilingual.AIME2025-Multilingual
Description
This repository contains a multi language version of the AIME2025 dataset.
As the english reference version, we haved used the one created by the authors of MathArena.
For completness, we have included the english version also in this repository, please, refer to the one contained in the MathArena github repository for the original one (https://github.com/eth-sri/matharena/tree/main/data/aime). Many thanks to Jasper Dekoninck for the help in understanding the structure… See the full description on the dataset page: https://huggingface.co/datasets/fedric95/AIME2025-Multilingual.aime26-multilingual
