datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mixed-address-parsing
mixed-address-parsing "📫"
Overview
The mixed-address-parsing dataset is designed to simulate the challenges encountered when processing real-world address inputs. It contains paired examples of noisy address strings (simulating user input) and their corresponding, clean, structured JSON responses. The dataset was generated by extracting components from open geocoding data and deliberately injecting multiple types of noise to mimic common human errors and input… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/mixed-address-parsing.proteinLM-mixed-pretraining-v1
Pretraining mix for Protein language models
In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources:
MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity.
UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.hindi-english-code-mixed-tweets-sentimentCode-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset
The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali.
Each record includes:
🆔 Id
🛒 ProductId
💬 Code-Mixed-Text
💡 Sentiment
The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication.
🌐 Text Distribution
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.mixed_hate_dataset
Mixed Harm–Safe Statements Dataset
WARNING: This paper contains potentially sensitive, harmful, and offensive content.
Paper | Code
Abstract
Recent progress in unsupervised probing methods — notably Contrast-Consistent Search (CCS) — has enabled the extraction of latent beliefs in language models without relying on token-level outputs.Since these probes offer lightweight diagnostic tools with low alignment tax, a central question arises:
Can they effectively… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/mixed_hate_dataset.Mixed-Math-Ins-Conversation-SplitCode_Mixed_video_Complaintleipzig_id_mixed_tufs4_2012_sentencesfinalized-mixed-dataset-v1nep-spell-mixed-evalMIT_mixed_final
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Thecoder3281f/MIT_mixed_final.mixeddataleipzig_id_mixed_tufs4_2012_wordprompt_sft_mixedgenerated_fitness_coach_mixedgraph_conn_mixed_id
