CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Peacockery /tajik-asr-corpus-v0 Tajik ASR Corpus v0 Deduplicated Tajik automatic speech recognition corpus assembled from FLEURS-derived speech data, Mozilla Common Voice 25 Tajik, and Muhtasham Tajik ASR augmented data. Format Each split has a data.tsv and an audio/ directory. TSV columns: id audio_filename raw_transcription transcription characters audio_bytes source source_id duplicate_count tajik_asr_combined.sqlite mirrors the TSV rows and includes normalized_text, source_split, and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v0.audioautomatic-speech-recognition1K<n<10K1 likes104 downloads4mo agoHugging Face02muhtasham /tajik-asr-augmented-testaudio1K<n<10K0 likes22 downloads1y agoHugging Face03shunyalabs /tajik-speech-datasetaudio1K<n<10K0 likes18 downloads1y agoHugging Face04muhtasham /tajik-asrgatedaudio10K<n<100K2 likes12 downloads1y agoHugging Face05Tohirju /tajik-voicechat-samplesgated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audion<1K1 likes9 downloads1mo agoHugging Face06Tohirju /tajik-voice-agent-demogated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audion<1K0 likes9 downloads1mo agoHugging Face07Tohirju /tajik-classic-audiobooksgated Tajik Classic Audiobooks Chapter-level audiobooks of classic Tajik/Persian literature, narrated by a single fine-tuned Chatterbox-multilingual Tajik TTS voice (synthetic). Text was cleaned to spoken form (numbers spelled out, no digits/Latin) and segmented per chapter. Each book folder holds mp3/ch_NNN.mp3 plus the source chapters_index.json and QC report. Books included: shakuri_khuroson, shakuri_panturkizm, sadi_guliston, sadi_buston, nizami_layli, nizami_makhzan_khusrav… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-classic-audiobooks.audiotext-to-speech1K<n<10K0 likes9 downloads1mo agoHugging Face08Tohirju /tajik-audiobooks-chaptersgated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audio10K<n<100K0 likes7 downloads1mo agoHugging Face09Tohirju /tajik-asrgated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audio10K<n<100K0 likes6 downloads1mo agoHugging Face10Tohirju /tajik-audiogated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audio10K<n<100K0 likes6 downloads1mo agoHugging Face11Tohirju /tajik-audiobooksgated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. audion<1K0 likes6 downloads1mo agoHugging Face12muhtasham /tajik-audiogatedaudio10K<n<100K5 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.