CoolFace
20 results

codeswitch

Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes287 downloads1y agoHugging FaceMoamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_new10K<n<100K0 likes230 downloads1y agoHugging Facegeorgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes196 downloads2y agoHugging Faceliva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes191 downloads2mo agoHugging Facecode-switching /text-summarizationtextsummarizationn<1K0 likes169 downloads24d agoHugging FacePraxel /codeswitch-pairs-lase-heldout Codeswitch Pairs LASE — Western held-out corpus 1043 held-out cross-script utterance pairs from 8 ElevenLabs Western Multilingual voices. Used to evaluate generalisation of speaker encoders trained on Praxel/codeswitch-pairs-lase. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-heldout.audioaudio-classification1K<n<10K0 likes161 downloads5mo agoHugging Face