datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.id-en-codeswitch-dataset-alternative
Indonesian–English Code-Switching Synthetic Speech Dataset
Synthetic speech generated for the undergraduate final project "Handling
Code-Switching in Automatic Speech Recognition for Low-Resource Language
Pairs: An Indonesian–English Case Study", School of Electrical Engineering
and Informatics, Institut Teknologi Bandung.
This dataset contains synthetic audio produced from the Indonesian–English
code-switching text corpora released in the companion repository below. It
was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.
