CoolFace
Datasetpublic

C1Tech/Persian-ASR-Benchmark

This dataset consists of 3 hours of 16kHz audio collected from diverse environments to better represent real-world scenarios. The recordings were sourced from audiobooks, YouTube, and other public sources, ensuring a wide variety of speech styles and acoustic conditions. One key advantage of this dataset is that it was collected from recent sources within the last few months, ensuring no overlap with training data and fairness for evaluating other STT models. To enable a robust and fair… See the full description on the dataset page: https://huggingface.co/datasets/C1Tech/Persian-ASR-Benchmark.

sourceHugging Faceupdated 2mo agoView on Hugging Face
4likes244downloads
Dataset Card

This dataset consists of 3 hours of 16kHz audio collected from diverse environments to better represent real-world scenarios. The recordings were sourced from audiobooks, YouTube, and other public sources, ensuring a wide variety of speech styles and acoustic conditions.

One key advantage of this dataset is that it was collected from recent sources within the last few months, ensuring no overlap with training data and fairness for evaluating other STT models.

To enable a robust and fair comparison of models, the dataset has been carefully normalized: extra characters and inconsistencies in Persian text have been removed, and orthographic variations have been standardized.

<p align="center"> <img src="assets/plot.png" alt="Evaluation of different models on the dataset" /> </p>

For a detailed explanation of the normalization process, please refer to our GitHub page.


We specialize in cutting-edge Voice & Audio Intelligence—from state-of-the-art Speech-to-Text (STT) and natural Text-to-Speech (TTS) to Voice Verification, Audio Intelligence, and domain-adapted LLMs. Need higher accuracy, lower latency, or custom-trained voice models for your enterprise? 📬 Contact Sales: info@c1tech.group 🔗 Explore Dashboard: https://ai.c1tech.group/dashboard