codeswitch
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_newcode_switch_yodas_zh
Dataset Card for code-switching yodas
This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas
This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon.
Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here.
Code-Switching ASR
Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.text-summarizationcodeswitch-pairs-lase-heldout
Codeswitch Pairs LASE — Western held-out corpus
1043 held-out cross-script utterance pairs from 8 ElevenLabs Western Multilingual voices. Used to evaluate generalisation of speaker encoders trained on Praxel/codeswitch-pairs-lase.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-heldout.
