datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-multidialect-emotional-speech-demo
DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech
A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request.
Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs.
gos-demo
Gronings transcribed speech
Demonstration dataset with Gronings transcribed speech based on the dataset released by San et al. (2021).
For more information see the corresponding ASRU 2021 paper.
nepal-oral-demo
Nepal Oral Demo
Public, always-safe fixtures for the Nepal oral-language backbone (AkAiNp).
Fictional “Demo Himalayan” track only
Schema examples for CI, export-script tests, and the Expo training app offline demo pack
No real community speakers, ever
Monorepo: nepal-multilingual-llm (local project). Source: data/packs/_demo/ + packages/schema/examples/.
Intended uses
OK
Not OK
Unit tests, pipeline dry-runs
Training production ASR/TTS as if it were… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-demo.
