CoolFace
Datasetpublic

yasalma/tat_hackathon_asr

Hackathon Tatar ASR Dataset Summary Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset… See the full description on the dataset page: https://huggingface.co/datasets/yasalma/tat_hackathon_asr.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
0likes79downloads
Dataset Card

Hackathon Tatar ASR

Dataset Summary

Hackathon Tatar ASR is a speech dataset distributed during the "Татар.Бу Хакатон" (Tatar.Bu Hackathon) held in Tatarstan in May 2024. This dataset likely consists of newly collected crowdsourced recordings created after the last release of TatSC (Tatar Speech Corpus), although some intersections with TatSC might be present. While TatSC contains 269.1 hours of transcribed speech with 271,914 utterances, this hackathon dataset comprises 89.99 hours of audio with 68,593 utterances.

Dataset Origin

The dataset was shared during the Tatar.Bu Hackathon 2024, which took place on May 17-19, 2024, at the IT Park in Kazan.

Dataset Structure

Parts: The dataset contains a single set of recordings, primarily comprised of crowdsourced audio samples.

Data Fields:

  • audio: The speech audio file in the Tatar language
  • text: The transcription of the audio file in Tatar
  • duration: Length of the audio in seconds
  • file_id: Unique identifier for the audio file
  • speaker_id: Identifier for the speaker in the recording

Data Processing

Audio Processing: The audio samples range from very short utterances (0.21s) to longer segments (31.1s).

Technical Details

Dataset TypeSpeech corpus for ASR
Languagett, Tatar
Speech StylePrimarily crowdsourced
ContentVaried
Audio Parameters16 kHz sampling rate
File FormatMP3, TXT (UTF-8)
Number of lines68,593
Total duration89.99 hours

Relation to Other Datasets

This dataset complements existing Tatar language resources such as the original TatSC (269.1 hours) and TatarTTS (~70 hours). It represents ongoing efforts to expand speech recognition capabilities for the Tatar language through community participation and hackathon initiatives.