datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Constrained_Indic_Codemixingko-en-code-mixing-sts
Korean–English Code-Mixing STS Dataset
This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred.
Interactive Dashboard
🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/
The dashboard provides:
Interactive data exploration and filtering
Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.fasttext_mixing_domains_top_2_codefasttext_mixing_domains_top_4_codefasttext_mixing_domains_top_8_codefasttext_mixing_domains_top_5_codefasttext_mixing_domains_top_6_codefasttext_mixing_domains_top_7_codefasttext_mixing_domains_top_3_codefasttext_mixing_domains_top_1_codeEnglish_French_safety_code-mixing_datasetUsing HarmBench promtps as the English baselines
code-mixing_safety
