CoolFace
Datasetpublic

dangrebenkin/long_audio_youtube_lectures

Audio dataset with Russian speech of scientific lectures from YouTube. The dataset contains seven long 20-40 minute Russian audios collected as a test set for our ASR system. The audios belong to different lexical and speech domains; they are parts of several Russian scientific lectures on various subjects: philology, mathematics, history, etc. All recordings were made in relatively quiet acoustic environments typical of lecture halls; however, some background noises, such as the sound of… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/long_audio_youtube_lectures.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
2likes63downloads
Dataset Card

Audio dataset with Russian speech of scientific lectures from YouTube.

The dataset contains seven long 20-40 minute Russian audios collected as a test set for our ASR system. The audios belong to different lexical and speech domains; they are parts of several Russian scientific lectures on various subjects: philology, mathematics, history, etc. All recordings were made in relatively quiet acoustic environments typical of lecture halls; however, some background noises, such as the sound of chalk hitting a blackboard, were present.

The dataset can be used as a test set for ASR systems, e.g. results for Pisets are described in https://aclanthology.org/2025.naacl-industry.74/.