CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01milanakdj /nepali-audio-reserve-r6gated Nepali two-speaker conversation chunks ~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.audioautomatic-speech-recognition10K<n<100K0 likes339 downloads15d agoHugging Face02milanakdj /nepali-tts-synthetic-v2gated Nepali TTS Synthetic v2 383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded as-is (no re-encode, no resampling). Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS synthesis → optional voice conversion against a pool of 600 real multi-speaker reference clips → ASR-based QC gate on character error rate. Read this before training on it Only 48% of rows are voice-converted. Each row carries a kept field recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.audiotext-to-speech100K<n<1M0 likes117 downloads24d agoHugging Face03milanakdj /nepali-audio-reserve-r5gated Nepali long-form single-speaker readings ~363.2 h of Nepali speech at 48 kHz. One speaker per clip, ~5 minute chunks. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party published readings whose rights holders did not grant redistribution. Access is gated and the source is not named. Columns column contents id first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r5.audioautomatic-speech-recognition1K<n<10K0 likes31 downloads16d agoHugging Face04milanakdj /nepali-audio-reserve-r1gated Nepali long-form speech, restored ~548.1 h of Nepali speech at 24 kHz. One speaker per clip, average 32 s, up to several minutes. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from AI4Bharat IndicVoices-R (CC-BY-4.0) and restored with sidon-v0.1. Attribution is required by that licence, so it is given here. These are the long-form files our quality gate rejected; the 74 h… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r1.audioautomatic-speech-recognition10K<n<100K0 likes31 downloads16d agoHugging Face05milanakdj /nepali-oov-distilledgated Nepali OOV-distilled subset (854 h) An OOV-dense distillation of Premal-12/c9nepali-audio-dataset2 (used with the author's permission), shipped in four variants: the original single-voice audio, a CPU-augmented copy, and 244 h re-rendered onto 1,842 real human speakers with Seed-VC. For Nepali ASR and TTS work. Filter with the variant field -- see Composition below. If you came here for speaker diversity, you want variant == "vc". What this is The source corpus is… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-oov-distilled.audioautomatic-speech-recognition10K<n<100K1 likes30 downloads13d agoHugging Face06milanakdj /nepali-audio-reserve-r2gated Nepali synthetic educational dialogue ~137.0 h of Nepali speech at 24 kHz. Two speakers per clip, 2-5 minutes, teacher/student turns. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Synthetic. Generated by a TTS model reading NCERT-style educational dialogue, code-mixed Nepali/English. No human speaker is recorded here. Columns column contents id first 16… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r2.audioautomatic-speech-recognition1K<n<10K0 likes19 downloads16d agoHugging Face07milanakdj /nepali-audio-reserve-r4gated Nepali short clips, ASR-agreed transcripts ~53.9 h of Nepali speech at 24 kHz. Single speaker, 3-20 s. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Transcripts kept only where three independent ASR systems agreed. Derived from third-party audio whose rights holders did not grant redistribution. Columns column contents id first 16 hex of sha256(original… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r4.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads16d agoHugging Face08milanakdj /nepali-audio-reserve-r3gated Nepali spontaneous short clips (SNR 40-50) ~16.2 h of Nepali speech at 24 kHz. Single speaker, 5-20 s, spontaneous. A backup, not a release: the transcripts are machine-generated, and none of this audio passed the quality gate that produced our training corpus. Derived from third-party web audio whose rights holders did not grant redistribution, which is why access is gated and the source is not named. Columns column contents id first 16 hex of… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r3.audioautomatic-speech-recognition1K<n<10K0 likes15 downloads16d agoHugging Face09fosters /knihi-be-arlou_milasc_kniazia_hieranima_output_original AudioSet Pipeline Output — арыгінальнае аўдыё Мова / Language: Беларуская (Belarusian) Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці. Частка калекцыі Ministerskija — корпус беларускіх аўдыёкніг. Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя): knihi-be-arlou_milasc_kniazia_hieranima_output Структура Кожны радок змяшчае: audio — арыгінальны аўдыёзапіс text — транскрыпцыя chunk_uid — унікальны ідэнтыфікатар Ліцэнзія /… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_milasc_kniazia_hieranima_output_original.audioautomatic-speech-recognitionn<1K0 likes13 downloads4mo agoHugging Face10fosters /knihi-be-arlou_milasc_kniazia_hieranima_all AudioSet Pipeline Output Мова / Language: Беларуская (Belarusian) Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд. Частка калекцыі Belarusian Audiobooks (native). Радкоў у датасеце 425 Працягласць 1 гадз 21 хв Частата дыскрэтызацыі 44100 Hz Каналы мона Даўжыня фрагмента да 30 с Структура Кожны радок змяшчае: audio — аўдыёфрагмент (native SR, мона, ≤30 с) text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-arlou_milasc_kniazia_hieranima_all.audioautomatic-speech-recognitionn<1K0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.