CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Josephgflowers /mixed-address-parsing mixed-address-parsing "📫" Overview The mixed-address-parsing dataset is designed to simulate the challenges encountered when processing real-world address inputs. It contains paired examples of noisy address strings (simulating user input) and their corresponding, clean, structured JSON responses. The dataset was generated by extracting components from open geocoding data and deliberately injecting multiple types of noise to mimic common human errors and input… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/mixed-address-parsing.text100K<n<1M0 likes521 downloads1y agoHugging Face02MichelNivard /proteinLM-mixed-pretraining-v1 Pretraining mix for Protein language models In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources: MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity. UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.texttoken-classification100M<n<1B0 likes354 downloads1y agoHugging Face03Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes84 downloads1y agoHugging Face04md-nishat-008 /Code-Mixed-Sentiment-Analysis-Dataset Dataset Generation: Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.text10K<n<100K0 likes59 downloads3y agoHugging Face05md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes39 downloads3y agoHugging Face06DaliaBarua /En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali. Each record includes: 🆔 Id 🛒 ProductId 💬 Code-Mixed-Text 💡 Sentiment The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication. 🌐 Text Distribution The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.text100K<n<1M0 likes25 downloads10mo agoHugging Face07SabrinaSadiekh /mixed_hate_dataset Mixed Harm–Safe Statements Dataset WARNING: This paper contains potentially sensitive, harmful, and offensive content. Paper | Code Abstract Recent progress in unsupervised probing methods — notably Contrast-Consistent Search (CCS) — has enabled the extraction of latent beliefs in language models without relying on token-level outputs.Since these probes offer lightweight diagnostic tools with low alignment tax, a central question arises: Can they effectively… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/mixed_hate_dataset.tabulartext-classification1K<n<10K5 likes25 downloads9mo agoHugging Face08prithivMLmods /Mixed-Math-Ins-Conversation-Splittext100K<n<1M2 likes19 downloads2y agoHugging Face09sudipghosh1728 /Code_Mixed_video_Complainttabulartext-classificationn<1K0 likes17 downloads2y agoHugging Face10nqzfaizal77ai /leipzig_id_mixed_tufs4_2012_sentencestext100K<n<1M0 likes16 downloads2y agoHugging Face11Yashodhar29 /finalized-mixed-dataset-v1text100K<n<1M0 likes10 downloads9mo agoHugging Face12dura-garage /nep-spell-mixed-evaltextn<1K0 likes9 downloads3y agoHugging Face13Thecoder3281f /MIT_mixed_finalgated Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Thecoder3281f/MIT_mixed_final.texttranslation1M<n<10M0 likes8 downloads10mo agoHugging Face14mx56748756 /mixeddatatextn<1K0 likes5 downloads3y agoHugging Face15nqzfaizal77ai /leipzig_id_mixed_tufs4_2012_wordtext100K<n<1M0 likes5 downloads2y agoHugging Face16xkwu /prompt_sft_mixedtext10K<n<100K0 likes4 downloads3y agoHugging Face17nicolasalt /generated_fitness_coach_mixedtext1K<n<10K0 likes2 downloads1y agoHugging Face18shivangchopra11 /graph_conn_mixed_idtabular1K<n<10K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.