CoolFace
11 results

code-mixing

Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes211 downloads1mo agoHugging Facewatchstep /ko-en-code-mixing-sts Korean–English Code-Mixing STS Dataset This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred. Interactive Dashboard 🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/ The dashboard provides: Interactive data exploration and filtering Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.tabularsentence-similarity1K<n<10K0 likes37 downloads1y agoHugging Facehafsamenaz1 /Group_H_Chinese-English-Code-Mixing-Prosody A Multimodal Dataset of Phonological Shift and Prosodic Reset in Chinese-English Code-Mixing Abstract Existing code-mixing corpora primarily rely on text transcripts, lacking the precise acoustic alignments necessary to study prosody at the switch boundary. This dataset provides 3 hours of carefully curated, naturalistic Chinese-English code-mixed speech sourced from diverse social media video content. We utilize a dual-level annotation scheme: manual token-level labeling… See the full description on the dataset page: https://huggingface.co/datasets/hafsamenaz1/Group_H_Chinese-English-Code-Mixing-Prosody.audion<1K1 likes24 downloads5mo agoHugging Facemlfoundations-dev /fasttext_mixing_domains_top_2_codetext1K<n<10K0 likes6 downloads2y agoHugging Facemlfoundations-dev /fasttext_mixing_domains_top_4_codetext1K<n<10K0 likes6 downloads2y agoHugging Facemlfoundations-dev /fasttext_mixing_domains_top_8_codetext1K<n<10K0 likes5 downloads2y agoHugging Face