datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.roman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.Roman-Urdu-Sentiment-Dataset
Roman Urdu Sentiment Dataset
This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis.
Dataset Structure
The dataset is formatted in comma-separated values (.csv) with the following columns:
text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.roman-urdu-sentiment-analysisroman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.roman-urdu-sentiment-embeddings
Roman Urdu Sentiment Embeddings Dataset
Overview
This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation.
Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.RomanUrdu-NLP-Sentiment-Corpus
RomanUrdu-NLP-Sentiment-Corpus
Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness
Overview
This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be:
Robust to slang and informal Roman Urdu
High-quality through LLM-assisted labeling and human validation
Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/RomanUrdu-NLP-Sentiment-Corpus.roman-urdu-alpaca-qa-mix
Dataset Card for Roman Urdu + Alpaca QA Mix
This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total:
500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API.
500 examples in English randomly sampled from the Stanford Alpaca dataset.
The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.Roman-Urdu-Toxic-Corpusmultilingual-urdu-romanurdu-arabic-english-sentiment
Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR)
Overview
This dataset is a clean and balanced multilingual sentiment classification dataset
covering four languages:
English
Urdu
Roman Urdu
Arabic
The dataset is designed to support sentiment analysis and text classification
tasks, especially for low-resource languages such as Urdu and Roman Urdu.
Sentiment Classes
Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.romanurdu-sentiment-datasetroman-urdu-corpus
Roman Urdu Corpus
Roman Urdu text collected from YouTube comments and Common Crawl (.pk
domains), filtered to keep likely Roman-Urdu lines and deduplicated.
roman_urdu_corpus.txt -- plain text, one entry per line
roman_urdu_corpus.jsonl -- same text with source metadata
Collected and filtered with a keyword-heuristic + dedup pipeline. No
guarantee of completeness or removal of all offensive/spam content --
inspect before using for training.
roman-urdu-blindspot-eval
Roman Urdu Blind Spot Evaluation — Qwen2.5-1.5B-Instruct
1. The Blind Spot
As a Pakistani speaker, I noticed that many people write everyday, harmless messages in Roman Urdu (Urdu written in English letters, mixed with English words) instead of pure English. This is how most Pakistanis actually communicate daily on WhatsApp and social media. I tested whether an AI model treats normal, harmless Roman Urdu messages differently than the same message in English.… See the full description on the dataset page: https://huggingface.co/datasets/muhammadhunzla971-sys/roman-urdu-blindspot-eval.roman_urdu_negation_blindspot
Roman Urdu Negation Blind Spot
Dataset Summary
This dataset tests whether Qwen3-4B-Instruct-2507 uses negation correctly when
judging Roman Urdu sentences, rather than just being able to translate it. It
pairs 12 sentences across two polarities (affirmative and negated), tested on
two tasks: English translation and a constrained Yes/No sentiment/factual
judgment, with a matched English control set (48 rows total). Built as the
technical challenge submission for… See the full description on the dataset page: https://huggingface.co/datasets/Tajwer/roman_urdu_negation_blindspot.Sentiment-Analysis-Roman-UrduLORU-ED_Roman-Urdu-Event-Detection
LORU-ED: Labeled Original Roman Urdu Dataset for Event Detection
Dataset Description
LORU-ED is a novel Roman Urdu dataset for event detection, containing 7,694
annotated sentences balanced equally between event and non-event classes.
Dataset Structure
Split
Size
Train
6,155
Test
1,539
Features
id: Unique identifier
text: Roman Urdu sentence
language: Language tag (roman_urdu/mixed/en)
is_event: Binary label (1=event, 0=non-event)… See the full description on the dataset page: https://huggingface.co/datasets/Adeel1percentdotcom/LORU-ED_Roman-Urdu-Event-Detection.roman_urduromanurdu_dataseturdu_roman_sentimentroman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original… See the full description on the dataset page: https://huggingface.co/datasets/hassanyra/roman_urdu.RomanUrdu-NLP-Sentiment-Corpus
RomanUrdu-NLP-Sentiment-Corpus
Largest Open-Source Roman Urdu Sentiment Dataset with Slang Robustness
Overview
This repository presents the largest publicly available Roman Urdu sentiment analysis dataset, containing 134,052 labeled text samples collected from chats and social media platforms. The dataset is designed to be:
Robust to slang and informal Roman Urdu
High-quality through LLM-assisted labeling and human validation
Balanced across sentiment classes… See the full description on the dataset page: https://huggingface.co/datasets/Inferencelab/RomanUrdu-NLP-Sentiment-Corpus.roman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original… See the full description on the dataset page: https://huggingface.co/datasets/MNajam/roman_urdu.romanurdu-sentiment-datasetParallel-Urdu-Roman-Urdu-CorpusRomanUrduEnglishRoman-Urdu-sentiment-1roman-urdu-retrieval-benchmarkroman-urdu-HateSpeech
