datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
ramanv-tts-codemixed-voiceCodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.Constrained_Indic_Codemixingcodemixed-id-hate-speech
Code-mixed Indonesian Hate Speech Dataset
Manually annotated hate speech dataset for Indonesian-Javanese and
Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations.
codemixCodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.codemixed-ind-classification
CodeMixed_ind_Classification
Deduplicated copy of kornwtp/codemixed-ind-classification.
Splits
split
rows
train
956
orca_dpo_pairs-Hinglish-Codemix
Summary
aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi
Version: 1.0
Citation
@misc {orca_dpo_pairs-Hinglish-Codemix,
author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.codemixqa
CodeMixQA
A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms.
We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.Code-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.code_mixed_jv_idSentiment analysis and machine translation data for Javanese and Indonesian.CodeMixBench
BigCodeBench-CodeMixed
Dataset Description
This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.nepali-english-codemixed-asrdatabricks-dolly-15k-Hinglish-Codemix
Summary
aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi
Version: 1.0
Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.codemixed-ind-classificationbelebele_hin_Latn_codemixedFiltered_Dakshina_Natural_CodeMixedcodemixed-ind-classificationcode-mixed-bangla-english-asrgujarati-english-codemixed-sentiment
Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset
Dataset Summary
This dataset addresses an empirically confirmed gap in the code-mixed
sentiment analysis literature: a systematic literature review of 48
retrieved records (39 core studies) found zero existing
Gujarati-English code-mixed sentiment datasets, despite Gujarati having
over 55 million native speakers. This dataset provides the first
Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.Codemix_tamil_english_test
Dataset Card for "Codemix_tamil_english_test"
More Information needed
Code_Mixed_video_Complainten-hi-codemixed-corpustelugu-qa-codemixed
Telugu QA Paraphrases
A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing.
Dataset Description
This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing.
Each example contains:
question : Original English question
answer : Ground-truth answer
level_0 : English paraphrase
level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.orca_dpo_pairs-Hinglish-Codemix
Summary
aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi
Version: 1.0
Citation
@misc {orca_dpo_pairs-Hinglish-Codemix,
author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/CharuAgarwal/orca_dpo_pairs-Hinglish-Codemix.nepali-english-codemixed-asrCodemix_tamil_englishArzEn-CodeMixed
ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset
This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants.
Dataset Details
Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.
