datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gujarati-english-codemixed-sentiment
Gujarati-English / Hindi-English Code-Mixed Sentiment Dataset
Dataset Summary
This dataset addresses an empirically confirmed gap in the code-mixed
sentiment analysis literature: a systematic literature review of 48
retrieved records (39 core studies) found zero existing
Gujarati-English code-mixed sentiment datasets, despite Gujarati having
over 55 million native speakers. This dataset provides the first
Gujarati-English resource of this kind, alongside… See the full description on the dataset page: https://huggingface.co/datasets/ShrutiPatel3011/gujarati-english-codemixed-sentiment.Code_Mixed_video_ComplaintArzEn-CodeMixed
ArzEn-CodeMixed: English-Arabic Intra-Sentential Code-Mixed Translation Dataset
This dataset contains aligned English, Arabic, and English-Arabic as well as Arabic-English intra-sentential code-mixed sentences, created for evaluating large LLMs on translating code-mixed text. It is constructed from the ArzEn-MultiGenre parallel corpus and enhanced with carefully prompted and human-validated code-mixed variants.
Dataset Details
Instances: Each entry contains:… See the full description on the dataset page: https://huggingface.co/datasets/taha-alnasser/ArzEn-CodeMixed.
