datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.CodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.telugu-qa-codemixed
Telugu QA Paraphrases
A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing.
Dataset Description
This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing.
Each example contains:
question : Original English question
answer : Ground-truth answer
level_0 : English paraphrase
level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.CodeMix
CodeMix - a small finetune dataset
6k chat/response pairs, a balanced mix of:
glaive-function-calling-v2 (agentic tool calling)
hermes-function-calling-v1 (tool calling)
CodeAlpaca-20k (coding)
dolly-15k (instruct)
marathi-codemix-qa
Marathi Minglish QA
~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles.
Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms.
Example
Question:
Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi?
Answer:
Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.
