datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EmiratiTTS-smoke-samples
EmiratiTTS — Stage 0.5 LoRA Smoke Samples
These 10 audio clips are the stage 0.5 acceptance check for the EmiratiTTS
project (Chatterbox Multilingual fine-tuned for Emirati Arabic).
This is NOT a model release. It is a sanity check that the data + tokenizer
reference-clip + ChatterboxMultilingualTTS pipeline is wired correctly before
committing GPUs to the long full-FT run. Quality is irrelevant at this stage —
the only pass criterion is "intelligible Arabic from both reference… See the full description on the dataset page: https://huggingface.co/datasets/Alqayed2024/EmiratiTTS-smoke-samples.alyah-emirati-benchmark
Alyah ⭐️: Emirati Dialect Benchmark for Arabic LLMs
📝 Blogpost |
🔧 Code on GitHub
Dataset Summary
Alyah (الياه in Emirati means North Star) is a manually curated evaluation benchmark for assessing the performance of Arabic large language models on the Emirati dialect. It is designed to measure not only linguistic competence, but also cultural awareness, pragmatic understanding, and sensitivity to social norms as expressed in everyday Emirati Arabic.
The… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/alyah-emirati-benchmark.EmiratiDialictShowsAudioTranscriptionThis dataset contains two files: a zipped file with segmented audio files from Emirati TV shows, podcasts, or YouTube channels, and a tsv file containing the transcription of the zipped audio files.
The purpose of the dataset is to act as a benchmark for Automatic Speech Recognition models that work with the Emirati dialect.
The dataset is made so that it covers different categories: traditions, cars, health, games, sports, and police.
Although the dataset is for the emirati dialect… See the full description on the dataset page: https://huggingface.co/datasets/eabayed/EmiratiDialictShowsAudioTranscription.alsanaa-emirati-arabic-asr
Alsanaa — Traditional Emirati Arabic Speech Dataset (ASR)
A curated and preprocessed corpus for traditional Emirati (Gulf) Arabic Automatic
Speech Recognition. It pairs short Emirati-dialect speech recordings with cleaned
Arabic transcriptions, covering customs, etiquette, and oral tradition.
Examples: 102 (train (102))
Approx. total audio: 4.54 hours
Audio format on the Hub: decoded Audio feature at 16000 Hz
Language: Arabic — traditional Emirati / Gulf dialect (ar)… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/alsanaa-emirati-arabic-asr.TTS_Female_Emiratidetails_airev-ai__emirati-14b-v2
Dataset Card for Evaluation run of airev-ai/emirati-14b-v2
Dataset automatically created during the evaluation run of model airev-ai/emirati-14b-v2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_airev-ai__emirati-14b-v2.details_airev-ai__emirati-14b-v3
Dataset Card for Evaluation run of airev-ai/emirati-14b-v3
Dataset automatically created during the evaluation run of model airev-ai/emirati-14b-v3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_airev-ai__emirati-14b-v3.EmiratiAudioDataset-20Hemiratiemiratitts-samples
