datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Til-Audio
Til-Audio
Қазақ тіліндегі аудио және транскрипттер · Казахское аудио с транскриптами · Kazakh audio with transcripts
Қазақша · Русский · English
Қазақша
Til-Audio — көлемі 249.77 ГБ, жүктелетін default конфигурациясында 380 068 аудио жазбасы мен транскрипті бар қазақ тіліндегі сөйлеу датасеті. Ол ASR, мәтіннен сөйлеу синтезі және аудио жіктеу тапсырмаларына арналған.
Құрамы мен құрылымы
Әр жолда audio пішіміндегі дыбыс, transcript мәтіні және lang… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-Audio.til26-asr-polyglot-lion-split
TIL 26 ASR Stratified Split
Source data was staged from /home/jupyter/advanced/asr without modifying the source folder.
The split is stratified by language with seed 3407 and validation ratio 0.2.
Files
train/train.jsonl with audio under train/audio/
val/val.jsonl with audio under val/audio/
Counts
Train: 3595
Val: 899
Class Counts
Split
Class
Count
train
chinese
900
train
english
897
train
malay
900
train
tamil… See the full description on the dataset page: https://huggingface.co/datasets/byumbyum/til26-asr-polyglot-lion-split.Til-Audio-Corpus-KK-v1
Til-Audio-Corpus-KK-v1
Қазақша аудио мен толық транскрипттер корпусы · Корпус казахского аудио с полными транскриптами · Kazakh audio corpus with full transcripts
Қазақша · Русский · English
Қазақша
Til-Audio-Corpus-KK-v1 — аудиокітаптар, «Шалқар» жинағы және мінбер уағыздарынан алынған қазақша дыбыс пен толық транскрипттер корпусы. Репозиторий көлемі — 53.99 ГБ; жинақта 3,303 аудио-мәтін жұбы және 1042.8 сағат жазба бар.
Құрамы
Дереккөз
Жұп… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-Audio-Corpus-KK-v1.til26-asr-split
TIL 26 ASR Stratified Split
Source data was staged from /home/jupyter/advanced/asr without modifying the source folder.
The split is stratified by language with seed 3407 and validation ratio 0.2.
Files
train/train.jsonl with audio under train/audio/
val/val.jsonl with audio under val/audio/
Counts
Train: 3595
Val: 899
Class Counts
Split
Class
Count
train
chinese
900
train
english
897
train
malay
900
train
tamil… See the full description on the dataset page: https://huggingface.co/datasets/albagon/til26-asr-split.
