datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Creativeguy-persian-asr
persian ASR Dataset
Audio resampled to 16000 Hz. Transcripts in Persian.
Splits
Split
Samples
Total Duration (H:MM:SS)
Volume
train
13165
100:37:38
2.712 GB
test
5643
43:08:16
1.163 GB
total
18808
143:45:55
3.874 GB
cretan-speech-corpusCretan is a variety of Modern Greek predominantly used by
speakers who reside on the island of Crete or belong to the Cretan
diaspora. This includes communities of Cretan origin that were
relocated to the village of Hamidieh in Syria and to Western
Asia Minor, following the population exchange between Greece
and Turkey in 1923. The historical and geographical factors
that have shaped the development and preservation of the dialect
include the long-term isolation of Crete from the mainland, and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/cretan-speech-corpus.zilora-haitian-creole-speech
Haitian Creole Natural Speech Corpus
Dataset Description
This dataset contains natural, spontaneous Haitian Creole speech from native speakers across multiple dialects.
Features
Audio files: Natural conversations, not read speech
Transcriptions: Manual, native-verified
Dialects: Port-au-Prince, Cap-Haïtien, Les Cayes, diaspora
Code-switching: Includes French and English code-switch instances
Citation
@misc{zilora2026creole,
title={Haitian Creole… See the full description on the dataset page: https://huggingface.co/datasets/ZiloraSystems/zilora-haitian-creole-speech.
