CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes236 downloads1y agoHugging Face02Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes86 downloads1y agoHugging Face03Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes77 downloads9mo agoHugging Face04md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes63 downloads3y agoHugging Face05md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face06kornwtp /codemixed-ind-classificationtextn<1K0 likes37 downloads2y agoHugging Face07DaliaBarua /En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali. Each record includes: 🆔 Id 🛒 ProductId 💬 Code-Mixed-Text 💡 Sentiment The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication. 🌐 Text Distribution The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.text100K<n<1M0 likes25 downloads10mo agoHugging Face08ShrutiPatel3011 /gujarati-english-codemixed-sentiment Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset Dataset Summary This dataset addresses an empirically confirmed gap in the code-mixed sentiment analysis literature: a systematic literature review of 48 retrieved records (39 core studies) found zero existing Gujarati-English code-mixed sentiment datasets, despite Gujarati having over 55 million native speakers. This dataset provides the first Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.tabulartext-classification10K<n<100K0 likes20 downloads1mo agoHugging Face09sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes17 downloads2y agoHugging Face10Tngarg /Codemix_tamil_englishtext10K<n<100K0 likes13 downloads3y agoHugging Face11taha-alnasser /ArzEn-CodeMixed ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants. Dataset Details Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.tabulartranslation1K<n<10K2 likes11 downloads1y agoHugging Face12karanverma19 /Advanced_CodeMix_Normalization_Dataset_India Evaluation & Benchmarking To validate dataset usefulness, normalization accuracy can be evaluated using: Exact Match Accuracy BLEU Score for text similarity Human evaluation for real-world correctness This dataset is designed to improve performance of multilingual NLP systems in handling noisy, code-mixed Indian queries. Data Transformation Approach The dataset was created by transforming real-world code-mixed queries into structured English. Variations include:… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/Advanced_CodeMix_Normalization_Dataset_India.textn<1K0 likes5 downloads6mo agoHugging Face13teppap /code-mix-thai-engtabular1K<n<10K0 likes2 downloads7mo agoHugging Face14karanverma19 /CodeMix_Query_Normalization_India CodeMix Query Normalization (India) Overview This dataset contains code-mixed user queries from Indian contexts, primarily in Hinglish and Punjabi, normalized into clean English. It reflects how users naturally communicate in real-world scenarios by mixing local languages with English. Features 100 high-quality samples Code-mixed queries (Hinglish, Punjabi) Clean normalized English outputs Real-world, informal user language patterns Covers domains such as… See the full description on the dataset page: https://huggingface.co/datasets/karanverma19/CodeMix_Query_Normalization_India.textn<1K0 likes2 downloads6mo agoHugging Face15chloeclerc17 /English_French_safety_code-mixing_datasetUsing HarmBench promtps as the English baselines textn<1K0 likes2 downloads5mo agoHugging Face16adealvii /codemixed-synthetic-sarc-11ktext10K<n<100K0 likes1 downloads1y agoHugging Face17chloeclerc17 /code-mixing_safetytextn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.