datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kazakh-stt
Kazakh Speech Dataset (KSD)
1. Dataset Summary
Purpose: High-quality, open-source Kazakh speech dataset for Automatic Speech Recognition (ASR) system development.
Developed by: Department of Artificial Intelligence and Big Data, Al-Farabi Kazakh National University.
Total Duration: 554 hours of recorded speech.
Total Number of Speakers: 873
Average Sentences per Speaker: 250 sentences (utterances).
Total Utterances: 204,250
File Format: .wav
Audio Characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/kazakh-stt.FLEURS-AR-EN-split
FLEURS-AR-EN Dataset
Dataset Description
FLEURS-AR-EN is an Arabic-to-English dataset designed for Speech Translation tasks. This dataset is derived from Google's FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) dataset, specifically focusing on aligned Arabic audio samples with their corresponding Arabic transcriptions and English translations.
Overview
Task: Speech Translation
Languages: Arabic (source) → English (target)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/farahabdou/FLEURS-AR-EN-split.farahan-yoruba-elder-corpus
Apakose Ezekiel Imoleayo — Farahàn
Appear. Be found.
Lagos, Nigeria · UNILAG Yoruba Studies ·
Graduating 2030
What I Build
Yoruba Oral Knowledge Corpus
I document primary-source Yoruba knowledge
directly from elder speakers in Lagos and
southwest Nigeria. Structured interviews
covering proverbs (Owe), oral history (Itan),
praise poetry (Oriki), and cultural knowledge
systems that no web scrape produces.
What makes this corpus different:… See the full description on the dataset page: https://huggingface.co/datasets/Apakose-Ezekiel/farahan-yoruba-elder-corpus.FLEURS-AR-ENVoice
