datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fa-makarem-kabiri-16kbps
Persian translation (Makarem) read by Kabiri
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
fa-makarem-kabiri-16kbps
Style
unknown
Riwayah
hafs
Kind
translation
Bitrate
16 kbps
Ayah files
6236 (235 MiB)
Verified against the upstream MD5 list
6236
Ayahs absent upstream
0
Upstream folder
translations/Makarem_Kabiri_16Kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/fa-makarem-kabiri-16kbps.Bangla-Youtube-audio-transcription-datasetcommon-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-huge.kabyle-piper-22khzkabyle_asrkab-noise-asxerxeckabuverdianu-speech-datasetkabyleKabyle_ASR-En_Translationcommon-voice-scripted-speech-kab-26-tiny
Common Voice Scripted Speech 26.0 - Kabyle (Cleaned)
This is a cleaned, speaker-disjoint subset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Step
Input
Output
Filter
Quality filter
609,940
573,073
≥2 upvotes, 0 downvotes
Character… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-tiny.kab_audioKabyle_ASR-Fr_Translationkabyle-audio-taqbaylit-languageOldDumpDatadataset_kabyle
Dataset Card for "dataset_kabyle"
More Information needed
tatoeba-kabyle-audio
Tatoeba Kabyle Audio Dataset
A clean, standardized audio-text dataset for Kabyle (Taqbaylit) automatic speech recognition, extracted from the Tatoeba Project and rigorously orthographically corrected.
Dataset Description
This dataset contains 47,789 Kabyle sentences with audio recordings (~25.78 hours total) sourced from Tatoeba. All transcriptions have been standardized to use correct Kabyle Latin characters, replacing visually similar false friends from Greek… See the full description on the dataset page: https://huggingface.co/datasets/boffire/tatoeba-kabyle-audio.EngASRwithCVWavKabyle_TTSVOZTORRESvaani-chhattisgarh_kabirdham-cleanedresearch-agent-briefingsviktar-karamazau-dzialba-kabanchyka
Дзяльба кабанчыка
Metadata
Author: Віктар Карамазаў
Title: Дзяльба кабанчыка
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/viktar-karamazau-dzialba-kabanchyka.Ov-yor
