datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PolyglotAudio
Citation
If you use this dataset in your research or downstream work, please cite:
@misc{polyglot_audio_2026,
author = {Fernandes, Reuben Chagas},
title = {PolyglotAudio: Multilingual Audio Pre-training Corpus},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/Reubencf/PolyglotAudio}}
}
APA-style:
Reuben Chagas Fernandes (2026). PolyglotAudio: Multilingual Audio Pre-training Corpus [Dataset]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/PolyglotAudio.polyglot-modelsThis dataset repository shall serve as a mirror hosting models for polyglot.
Availability
Please note that currently the only languages, which have all models, are English (en) and Bulgarian (bg). Other languages may have partial support.
Adding missing models
In case you have previously downloaded polyglot language models, which are not available in this repo, please open a Pull Request.
License
All rights belong to the original authors. Please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/ndandanov/polyglot-models.til26-asr-polyglot-lion-split
TIL 26 ASR Stratified Split
Source data was staged from /home/jupyter/advanced/asr without modifying the source folder.
The split is stratified by language with seed 3407 and validation ratio 0.2.
Files
train/train.jsonl with audio under train/audio/
val/val.jsonl with audio under val/audio/
Counts
Train: 3595
Val: 899
Class Counts
Split
Class
Count
train
chinese
900
train
english
897
train
malay
900
train
tamil… See the full description on the dataset page: https://huggingface.co/datasets/byumbyum/til26-asr-polyglot-lion-split.polyglot-audio-samples
