datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
antalia-voice-corpus
Antalia Turkish single-speaker scripted speech
5.008 hours / 1,073 clips of studio-quality Turkish read speech from one professional voice
actor, script-aligned, segmented, and quality-gated. This is the corpus the
Antalia 1 voice was fine-tuned on.
Clean, consented, single-speaker Turkish speech at this quality is scarce — which is the main
reason this release exists. Development of the model is discontinued; the data is published
as-is so it stays useful.
Model:… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-voice-corpus.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.
