datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.east-african-code-switching
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
east_african_code_switching
This dataset contains utterances demonstrating code-switching between English and various East African languages, including Swahili, Luganda, and Acholi. Each sample is annotated with metadata such as region, intent, domain, sentiment, and specific switch types to support linguistic analysis. The collection covers diverse conversational registers… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/east-african-code-switching.
