datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ko-en-code-mixing-sts
Korean–English Code-Mixing STS Dataset
This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred.
Interactive Dashboard
🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/
The dashboard provides:
Interactive data exploration and filtering
Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.hero_run_3_code_mix_all_shardscode_rlvr_mixture_dpocode_rlvr_mixture_dpogujarati-english-codemixed-sentiment
Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset
Dataset Summary
This dataset addresses an empirically confirmed gap in the code-mixed
sentiment analysis literature: a systematic literature review of 48
retrieved records (39 core studies) found zero existing
Gujarati-English code-mixed sentiment datasets, despite Gujarati having
over 55 million native speakers. This dataset provides the first
Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.Code_Mixed_video_Complaintcode_rlvr_mixture_sftArzEn-CodeMixed
ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset
This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants.
Dataset Details
Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.marathi-codemix-qa
Marathi Minglish QA
~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles.
Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms.
Example
Question:
Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi?
Answer:
Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.train_fasttext_classifier_seed_code_best_mixtrain_fasttext_classifier_seed_code_worst_mix_3_1train_fasttext_classifier_seed_code_worst_mix_8_5_3_1train_fasttext_classifier_seed_code_worst_mix_5_3_1train_fasttext_classifier_seed_code_worst_mix_12_8_5_3_1code-mix-thai-eng
