CoolFace
Datasetpublic

lab260/espeech_balalaika

ESpeech datasets (w/o podcasts) Annotated by Balalaika [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/espeech_balalaika.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
4likes175downloads
filebuldjat_stripped_balalaika.parquet6.1 MBdownload
filebuldjat_stripped.tar.gz1.82 GBdownload
fileigm_balalaika.parquet46.3 MBdownload
fileigm.tar.gz23.73 GBdownload
filetuchniyzhab_stripped_balalaika.parquet47.0 MBdownload
filetuchniyzhab_stripped.tar.gz197.1 MBdownload
filewebinars_stripped_balalaika.parquet33.1 MBdownload
filewebinars_stripped.tar.gz11.75 GBdownload

lab260/espeech_balalaika · main · files are served by the source, never re-hosted here