datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.Roman-Urdu-Sentiment-Dataset
Roman Urdu Sentiment Dataset
This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis.
Dataset Structure
The dataset is formatted in comma-separated values (.csv) with the following columns:
text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.roman-urdu-sentiment-embeddings
Roman Urdu Sentiment Embeddings Dataset
Overview
This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation.
Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.RomanUrdu-NLP-Sentiment-Corpus
RomanUrdu-NLP-Sentiment-Corpus
Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness
Overview
This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be:
Robust to slang and informal Roman Urdu
High-quality through LLM-assisted labeling and human validation
Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/RomanUrdu-NLP-Sentiment-Corpus.multilingual-urdu-romanurdu-arabic-english-sentiment
Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR)
Overview
This dataset is a clean and balanced multilingual sentiment classification dataset
covering four languages:
English
Urdu
Roman Urdu
Arabic
The dataset is designed to support sentiment analysis and text classification
tasks, especially for low-resource languages such as Urdu and Roman Urdu.
Sentiment Classes
Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.roman_urdu_negation_blindspot
Roman Urdu Negation Blind Spot
Dataset Summary
This dataset tests whether Qwen3-4B-Instruct-2507 uses negation correctly when
judging Roman Urdu sentences, rather than just being able to translate it. It
pairs 12 sentences across two polarities (affirmative and negated), tested on
two tasks: English translation and a constrained Yes/No sentiment/factual
judgment, with a matched English control set (48 rows total). Built as the
technical challenge submission for… See the full description on the dataset page: https://huggingface.co/datasets/Tajwer/roman_urdu_negation_blindspot.Sentiment-Analysis-Roman-Urduromanurdu_datasetRomanUrdu-NLP-Sentiment-Corpus
RomanUrdu-NLP-Sentiment-Corpus
Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness
Overview
This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be:
Robust to slang and informal Roman Urdu
High-quality through LLM-assisted labeling and human validation
Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Inferencelab/RomanUrdu-NLP-Sentiment-Corpus.Roman-urdu_reviews_and_summaryRomanUrdu_Datasetromanurdupoetryurdu_Roman_Urdu_Addresses_Dataset
