CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RidheshBhati /Codemixed_New Codemixed ASR Dataset Unified collection of code-mixed ASR datasets. audioautomatic-speech-recognition100K<n<1M2 likes5.2k downloads5mo agoHugging Face02lingamvamshikrishnareddy /ramanv-tts-codemixed-voicegated0 likes517 downloads1mo agoHugging Face03CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes240 downloads1y agoHugging Face04Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes211 downloads1mo agoHugging Face05aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes199 downloads4mo agoHugging Face06mrzoadic /codemixaudio10K<n<100K0 likes85 downloads10mo agoHugging Face07Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes78 downloads9mo agoHugging Face08suwaimyo /codemixed-ind-classification CodeMixed_ind_Classification Deduplicated copy of kornwtp/codemixed-ind-classification. Splits split rows train 956 textn<1K0 likes77 downloads25d agoHugging Face09aaditya /orca_dpo_pairs-Hinglish-Codemix Summary aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Citation @misc {orca_dpo_pairs-Hinglish-Codemix, author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/orca_dpo_pairs-Hinglish-Codemix.text10K<n<100K1 likes64 downloads3y agoHugging Face10gentaiscool /codemixqa CodeMixQA A benchmark with high-quality human annotations, comprising 16 diverse parallel code-switched language-pair variants that span multiple geographic regions and code-switching patterns, and include both original scripts and their transliterated forms. We use SimpleQA Verified as our source dataset. We select the SimpleQA Verified, as it is a challenging evaluation set that has not been saturated yet by current models and has desirable properties such as verifiable answers… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/codemixqa.text10K<n<100K1 likes62 downloads8mo agoHugging Face11md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes58 downloads3y agoHugging Face12SEACrowd /code_mixed_jv_idSentiment analysis and machine translation data for Javanese and Indonesian.0 likes57 downloads2y agoHugging Face13ColdSlim /CodeMixBench BigCodeBench-CodeMixed Dataset Description This dataset is an augmented version of BigCodeBench designed for evaluating code generation in code-mixed scenarios. It introduces multilingual variations of the prompts, primarily focusing on translating the docstrings within the complete_prompt field while keeping the code and test cases in English. This allows for assessing the ability of Large Language Models (LLMs) to handle code generation tasks where the prompt contains a… See the full description on the dataset page: https://huggingface.co/datasets/ColdSlim/CodeMixBench.text1K<n<10K0 likes52 downloads1y agoHugging Face14Shyyamsh /nepali-english-codemixed-asraudio100K<n<1M0 likes52 downloads4mo agoHugging Face15aaditya /databricks-dolly-15k-Hinglish-Codemix Summary aaditya/databricks-dolly-15k-Hindi is an open source Hinglish-Codemix version dataset of databricks/databricks-dolly-15k. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Original Dataset repo… See the full description on the dataset page: https://huggingface.co/datasets/aaditya/databricks-dolly-15k-Hinglish-Codemix.text10K<n<100K2 likes51 downloads3y agoHugging Face16md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes38 downloads3y agoHugging Face17kornwtp /codemixed-ind-classificationtextn<1K0 likes38 downloads2y agoHugging Face18Thanmay /belebele_hin_Latn_codemixedtextn<1K0 likes35 downloads21d agoHugging Face19user-anto /Filtered_Dakshina_Natural_CodeMixedtext1K<n<10K0 likes31 downloads2mo agoHugging Face20puttatidam /codemixed-ind-classificationtextn<1K0 likes30 downloads7d agoHugging Face21fayez94 /code-mixed-bangla-english-asraudio1K<n<10K0 likes28 downloads2y agoHugging Face22ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes20 downloads1mo agoHugging Face23Tngarg /Codemix_tamil_english_test Dataset Card for "Codemix_tamil_english_test" More Information needed textn<1K0 likes18 downloads3y agoHugging Face24sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes17 downloads2y agoHugging Face25HuggMachas /en-hi-codemixed-corpustext1K<n<10K0 likes16 downloads2y agoHugging Face26nlpctx /telugu-qa-codemixed Telugu QA Paraphrases A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing. Dataset Description This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing. Each example contains: question : Original English question answer : Ground-truth answer level_0 : English paraphrase level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.textquestion-answering1K<n<10K0 likes16 downloads3mo agoHugging Face27CharuAgarwal /orca_dpo_pairs-Hinglish-Codemix Summary aaditya/orca_dpo_pairs-Hinglish-Codemix is an open source Hinglish version dataset of Intel/orca_dpo_pairs This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Citation @misc {orca_dpo_pairs-Hinglish-Codemix, author = { Pal, Ankit }… See the full description on the dataset page: https://huggingface.co/datasets/CharuAgarwal/orca_dpo_pairs-Hinglish-Codemix.text10K<n<100K0 likes14 downloads9mo agoHugging Face28kkarhm /nepali-english-codemixed-asraudio100K<n<1M0 likes14 downloads6mo agoHugging Face29Tngarg /Codemix_tamil_englishtext10K<n<100K0 likes13 downloads3y agoHugging Face30taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.