CoolFace
Datasetpublic

dangrebenkin/sova_rudevices_audiobooks

Dataset instance structure {'audio': {'path': '/path/to/wav.wav', 'array': array([wav numpy array]), dtype=float32), 'sampling_rate': 16000}, 'transcription': 'транскрипция'} Dataset audio info 16000 Hz wav mono Russian speech from audiobooks Citation @misc{sova2021rudevices, author = {Zubarev, Egor and Moskalets, Timofey and SOVA.ai}, title = {SOVA RuDevices Dataset: free public STT/ASR dataset with manually annotated live speech}… See the full description on the dataset page: https://huggingface.co/datasets/dangrebenkin/sova_rudevices_audiobooks.

sourceHugging Faceapache-2.0updated 4y agoView on Hugging Face
2likes714downloads
Dataset Card

Dataset instance structure

{'audio': {'path': '/path/to/wav.wav', 'array': array([wav numpy array]), dtype=float32), 'sampling_rate': 16000}, 'transcription': 'транскрипция'}

Dataset audio info

  • —16000 Hz
  • —wav
  • —mono
  • —Russian speech from audiobooks

Citation

@misc{sova2021rudevices, author = {Zubarev, Egor and Moskalets, Timofey and SOVA.ai}, title = {SOVA RuDevices Dataset: free public STT/ASR dataset with manually annotated live speech}, publisher = {GitHub}, journal = {GitHub repository}, year = {2021}, howpublished = {\url{https://github.com/sovaai/sova-dataset}}, }