datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KambaBench-ASR
KambaBench-ASR
Status: v0.0 — scaffold. No evaluation audio or gold transcriptions have been finalized yet.
An open, leakage-controlled, reproducible evaluation benchmark for Kamba (Kikamba, kam) automatic speech recognition (ASR).
KambaBench-ASR is designed to provide a common evaluation standard for Kamba speech-recognition systems. The benchmark is intended to be model-agnostic: any Kamba ASR system, whether based on Whisper, MMS, Omnilingual ASR, Parakeet, or another… See the full description on the dataset page: https://huggingface.co/datasets/lawmaluki/KambaBench-ASR.tajik-law-audio
Tajik Law Audio (1990–2025)
Synthetic Tajik speech read from the laws of the Republic of Tajikistan.
138,994 clips · ~413 hours · 2 voices (male + female)
Covers laws adopted 1990–2025, plus conventions, amnesty acts and the Constitution
Article structure preserved: law · year · chapter · article · clause
Numbers, dates and legal citations expanded to spoken Tajik
Round-trip QC against a Tajik ASR model: mean CER 0.051, median 0.020
Contents
file
what… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-law-audio.
