datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.SYNTH-Swallow-Math-Code-Mix
Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2
This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources:
SYNTH ~63.5%
SwallowCode-v2 ~15.5%
SwallowMath-v2-textbook ~10.5%
SwallowMath-v2-qa ~10.0%
The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.ramanv-tts-codemixed-voiceSinhala-English-Code-Mixed-Code-Switched-Dataset
Sinhala-English-Code-Mixed-Code-Switched-Dataset
This dataset contains 10,000 comments that have been annotated at the sentence level for sentiment analysis, humor detection, hate speech detection, aspect identification, and language identification.
The following is the tag scheme.
Sentiment - Positive, Negative, Neutral, Conflict
Humor - Humorous, Non humorous
Hate Speech - Hate-Inducing, Abusive, Not offensive
Aspect - Network, Billing or Price, Package, Customer Service, Data… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-English-Code-Mixed-Code-Switched-Dataset.CodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.codemixed-id-hate-speech
Code-mixed Indonesian Hate Speech Dataset
Manually annotated hate speech dataset for Indonesian-Javanese and
Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations.
Constrained_Indic_Codemixingmopd-math-code-mix
MOPD math+code mix
`train/`: math:code ≈ 1:1 平衡集(`math.parquet` + `code_*.parquet` shards)
`val/mopd_val_mix.parquet`: AIME24 全量 + MATH-500 子集 + Eurus code_validation 子集
路由字段:`ability ∈ {math, code}`
hindi-english-code-mixed-tweets-sentimentCodeMixBench
ℹ️Dataset Card for CodeMixBench
[EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation.
To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.codemixed-ind-classification
CodeMixed_ind_Classification
Deduplicated copy of kornwtp/codemixed-ind-classification.
Splits
split
rows
train
956
codemixmix60k-math-code-sft
mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture
This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.
The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base
main triad in the dLLM Registers project.
Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct.
License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.RLHFlow_mixture_clean_empty_round_with_dart_code_v1
Dataset Card for "RLHFlow_mixture_clean_empty_round_with_dart_code_v1"
More Information needed
orca_dpo_pairs-Hinglish-Codemix
Summary
aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi
Version: 1.0
Citation
@misc {orca_dpo_pairs-Hinglish-Codemix,
author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.code-trainer-v9-mixed
code-trainer-v9-mixed
40,401-row mixed training dataset for supervised fine-tuning (SFT) in the
Code-Trainer / RTPI pipeline.
Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training.
Composition
Slice
Source
Rows (train)
Purpose
A -- Code generation
cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample)
7,074
Preserve code-gen quality
B -- Tool calling
glaiveai/glaive-function-calling-v2 (19K cap)
~15,125
High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.codemixqa
CodeMixQA
A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms.
We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.Code-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.code_mixed_jv_idSentiment analysis and machine translation data for Javanese and Indonesian.hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below:
https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file
https://github.com/piyushmakhija5/hinglishNorm
https://github.com/ishan00/translation-for-code-switching-acl/tree/master
databricks-dolly-15k-Hinglish-Codemix
Summary
aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi
Version: 1.0
Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.nepali-english-codemixed-asrHINMIX_hi-en-code-mix-part-1CodeMixBench
BigCodeBench-CodeMixed
Dataset Description
This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.preference_dataset_mixture2_and_safe_pku30k_and_argilla_math_and_ultra_code_for_preference_modelcode-instruct-mixed
Description
Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset
ds = load_dataset("PotatoHD/code-instruct-mixed")
Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.codemixed-ind-classification
