CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.4k downloads5mo agoHugging Face02TMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes736 downloads8mo agoHugging Face03CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes234 downloads1y agoHugging Face04aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes209 downloads4mo agoHugging Face05Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes190 downloads1mo agoHugging Face06GXLXY /mopd-math-code-mix MOPD math+code mix `train/`: math:code ≈ 1:1 平衡集(`math.parquet` + `code_*.parquet` shards) `val/mopd_val_mix.parquet`: AIME24 全量 + MATH-500 子集 + Eurus code_validation 子集 路由字段:`ability ∈ {math, code}` texttext-generation10K<n<100K0 likes132 downloads1mo agoHugging Face07Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes84 downloads1y agoHugging Face08Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes78 downloads9mo agoHugging Face09suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes77 downloads25d agoHugging Face10mrzoadic /codemixaudio10K<n<100K0 likes75 downloads10mo agoHugging Face11albertge /mix60k-math-code-sft mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models. The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base main triad in the dLLM Registers project. Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct. License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.texttext-generation10K<n<100K0 likes70 downloads5d agoHugging Face12pxyyy /RLHFlow_mixture_clean_empty_round_with_dart_code_v1 Dataset Card for "RLHFlow_mixture_clean_empty_round_with_dart_code_v1" More Information needed text1M<n<10M0 likes68 downloads2y agoHugging Face13aaditya /orca_dpo_pairs-Hinglish-Codemix Summary aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Citation @misc {orca_dpo_pairs-Hinglish-Codemix, author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.text10K<n<100K1 likes66 downloads3y agoHugging Face14AmnaHassan /Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam Unity Code and GPT-Generated GDD Pairs Dataset This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications. Format Each entry is stored as a .jsonl file with: "input": GPT-4 generated GDD describing a specific game and its mechanics "output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.textn<1K4 likes65 downloads1y agoHugging Face15atlas-institute /code-trainer-v9-mixed code-trainer-v9-mixed 40,401-row mixed training dataset for supervised fine-tuning (SFT) in the Code-Trainer / RTPI pipeline. Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training. Composition Slice Source Rows (train) Purpose A -- Code generation cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample) 7,074 Preserve code-gen quality B -- Tool calling glaiveai/glaive-function-calling-v2 (19K cap) ~15,125 High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.text10K<n<100K0 likes62 downloads4d agoHugging Face16gentaiscool /codemixqa CodeMixQA A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms. We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.text10K<n<100K1 likes61 downloads8mo agoHugging Face17md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes59 downloads3y agoHugging Face18ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes56 downloads2y agoHugging Face19aaditya /databricks-dolly-15k-Hinglish-Codemix Summary aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.text10K<n<100K2 likes53 downloads3y agoHugging Face20Shyyamsh /nepali-english-codemixed-asraudio100K<n<1M0 likes52 downloads4mo agoHugging Face21upperwal /HINMIX_hi-en-code-mix-part-1audio10K<n<100K0 likes51 downloads1y agoHugging Face22ColdSlim /CodeMixBench BigCodeBench-CodeMixed Dataset Description This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.text1K<n<10K0 likes49 downloads1y agoHugging Face23weqweasdas /preference_dataset_mixture2_and_safe_pku30k_and_argilla_math_and_ultra_code_for_preference_modeltext100K<n<1M0 likes47 downloads2y agoHugging Face24PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face25md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face26kornwtp /codemixed-ind-classificationtextn<1K0 likes38 downloads2y agoHugging Face27scottgeng00 /olmo-3-preference-mix-deltas-complement2-yolo_victoria_hates_code-DECONtext100K<n<1M0 likes38 downloads1y agoHugging Face28watchstep /ko-en-code-mixing-sts Korean–English Code-Mixing STS Dataset This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred. Interactive Dashboard 🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/ The dashboard provides: Interactive data exploration and filtering Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.tabularsentence-similarity1K<n<10K0 likes37 downloads1y agoHugging Face29atlas-institute /code-trainer-v8-mixedtext10K<n<100K0 likes35 downloads2mo agoHugging Face30Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes35 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.