roman-urdu
roman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.roman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.Roman-Urdu-Sentiment-Dataset
Roman Urdu Sentiment Dataset
This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis.
Dataset Structure
The dataset is formatted in comma-separated values (.csv) with the following columns:
text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.roman-urdu-sentiment-analysisroman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.
