datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.Urdu-Multi-Domain-Benchmark
Urdu Multi-Domain Datasets
33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time.
Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore.
Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts.
Permanent archive: Zenodo DOI 10.5281/zenodo.22195610.
How to load
Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.ouhd-l-online-urdu-nastaliq-handwriting
OUHD-L: Online Urdu Nastaliq Handwriting — Line Pen Trajectories
Unmodified mirror. This repository re-hosts the OUHD-L v1.0 core release
exactly as published on Zenodo, byte-for-byte. Nothing has been added to or
removed from the data. It exists only to provide an alternative download
endpoint. The canonical source and citation is the Zenodo record:
https://zenodo.org/records/20642162 — DOI
10.5281/zenodo.20642162, version 1.0.0.
Overview
2,403 handwritten Urdu… See the full description on the dataset page: https://huggingface.co/datasets/saad2002/ouhd-l-online-urdu-nastaliq-handwriting.urdu-spam-dataset
Urdu Spam Detection Dataset
Description
This dataset is designed for classifying Urdu text into:
0 → Not Spam
1 → Spam
It is intended for AI-powered emergency helpline systems (e.g., 1122/911) to filter prank or irrelevant calls.
Dataset Structure
Format: CSV
Column
Type
Description
text
string
Urdu sentence
label
int (0/1)
Spam classification
Example
text,label
آپ کو میں نے پہلے بھی کال کیا تھا کیا یاد ہے,1
یہ… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-spam-dataset.Urdu_DataThis repo has cleaned urdu data scraped from the web.
urdu-idioms-with-english-translationurdu_rag_dataset.csv
Dataset Card for Urdu RAG Knowledge Base
Dataset Overview
This dataset is designed specifically to bootstrap and evaluate Retrieval-Augmented Generation (RAG) applications, search systems, and semantic retrieval pipelines using the Urdu language. It contains 185 clean, structured, and informative text chunks covering a wide array of domains.
Language: Urdu (ur)
Script: Nastaliq / Arabic script (Unicode UTF-8)
Total Rows: 185 chunks
Format: CSV (id, title… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/urdu_rag_dataset.csv.Urdu-Poetry-Dataset
Urdu Poetry Dataset
Welcome to the Urdu Poetry Dataset! This dataset is a collection of Urdu poems where each entry includes two columns:
Title: The title of the poem.
Poem: The full text of the poem.
This dataset is ideal for natural language processing tasks such as text generation, language modeling, or cultural studies in computational linguistics.
Dataset Overview
The Urdu Poetry Dataset comprises a diverse collection of Urdu poems, capturing the rich heritage… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-Poetry-Dataset.Roman-Urdu-Sentiment-Dataset
Roman Urdu Sentiment Dataset
This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis.
Dataset Structure
The dataset is formatted in comma-separated values (.csv) with the following columns:
text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.Urdu_Sentement
Roman Urdu Sentiment Analysis Dataset
A clean, structured dataset for sentiment analysis in Roman Urdu, split into training and testing sets. This dataset is optimized for quick integration with the Hugging Face datasets library and is ideal for fine-tuning text classification models (such as DistilBERT, mBERT, or XLM-RoBERTa) to handle Roman Urdu customer feedback, reviews, and social media text.
Dataset Structure
The dataset contains short user reviews and text… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Urdu_Sentement.roman-urdu-sentiment-embeddings
Roman Urdu Sentiment Embeddings Dataset
Overview
This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation.
Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.Roman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.multilingual-urdu-romanurdu-arabic-english-sentiment
Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR)
Overview
This dataset is a clean and balanced multilingual sentiment classification dataset
covering four languages:
English
Urdu
Roman Urdu
Arabic
The dataset is designed to support sentiment analysis and text classification
tasks, especially for low-resource languages such as Urdu and Roman Urdu.
Sentiment Classes
Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.Poly-FEVER-Urdu-Translation
Poly-FEVER Urdu Translation
Dataset Description
This dataset is an Urdu language extension of the original
Poly-FEVER
benchmark — a multilingual fact verification dataset for hallucination
detection in Large Language Models.
Urdu (اردو) is spoken by over 230 million people worldwide but was not
included in the original Poly-FEVER dataset, which covers 11 languages.
This dataset fills that gap by providing complete translations of all
77,971 factual claims… See the full description on the dataset page: https://huggingface.co/datasets/Urwashanza/Poly-FEVER-Urdu-Translation.Urdu-Turn-Detection-10k
Urdu Turn Detection Dataset 🗣️
A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete).
Dataset Details
Total Samples: 10,000
Language: Urdu (ur) - Nastaliq/Arabic Script only.
Cleanliness: - 100% Urdu Script (No Roman/English).
Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.Urdu-Multimodal-Emotion-Dataseturdu-binary-classification-dataThis Urdu sentiment dataset was formed by concatenating the following two datasets:
https://github.com/MuhammadYaseenKhan/Urdu-Sentiment-Corpus
https://www.kaggle.com/datasets/akkefa/imdb-dataset-of-50k-movie-translated-urdu-reviews
Sentiment-Analysis-Roman-Urduenglish-urduurdu-wordsEnglish-Urdu-Dataset
English Transliteration (Roman Urdu) of Urdu Poetry Dataset
Welcome to the English Transliteration (Roman Urdu) of Urdu Poetry Dataset! This dataset provides a collection of Urdu poetry transliterated into Roman script, making the rich literary heritage of Urdu accessible to a broader audience. Each entry in the dataset includes two columns:
Title: The transliterated title of the poem in Roman Urdu.
Poem: The full text of the poem in Roman Urdu.
This dataset is ideal for… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/English-Urdu-Dataset.UrduMultiDomainClassification
Urdu Multi-Domain Text Classification Dataset
Dataset Summary
This dataset is a multi-domain Urdu text classification dataset designed for sentiment analysis, intent recognition, topic classification, and binary relevance detection.It contains short Urdu sentences covering multiple real-world domains such as health, education, population, and general/other topics.
Each example is annotated with four labels:
Sentiment → positive, negative, neutral
Topic → health… See the full description on the dataset page: https://huggingface.co/datasets/umar178/UrduMultiDomainClassification.urdu_sa
Sentiment Analysis Data for the Urdu Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Khan et al. (2020).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{khan2017harnessing,
title={Harnessing English Sentiment Lexicons for Polarity Detection in Urdu Tweets: A Baseline Approach},
author={Khan, Muhammad Yaseen and Emaduddin, Shah Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/urdu_sa.Sentence_wise_urdu_text_dataset
Sentence_wise_urdu_text_dataset
Dataset Overview
File Information
Size: 5.29 MB (5,545,229 bytes)
Encoding: UTF-8
Basic Statistics
Total Characters: 3,136,348
Total Characters (excluding spaces): 2,472,408
Total Lines: 69,743
Total Words: 666,907
Linguistic Analysis
Vocabulary Size: 29,888
Average Word Length: 3.56 characters
Median Word Length: 3 characters
Average Paragraph Length: 670091.00 words
Hapax Legomena… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Sentence_wise_urdu_text_dataset.metanova_testUrdu-News
[Your Dataset Name]
Dataset Description
This dataset appears to be a collection of news headlines and their corresponding news text. Based on the provided image sample, the text content is in a language that uses the Arabic/Persian script, likely Persian (Farsi) or a similar Middle Eastern language. The dataset is structured in a tabular format suitable for various natural language processing tasks related to news content.
Dataset Structure
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-News.xnli2.0_train_urdulanguage: ["Urdu"]
urdu_ds2300urdfstudio:)
UrduG2P
Zuhri — Urdu G2P Dataset
Zuhri is a comprehensive and manually verified Urdu Grapheme-to-Phoneme (G2P) dataset. It is designed to aid research and development in areas such as speech synthesis, pronunciation modeling, and computational linguistics, specifically for the Urdu language.
This dataset provides accurate phoneme transcriptions and IPA representations, making it ideal for use in building high-quality TTS (Text-to-Speech), ASR (Automatic Speech Recognition), and other… See the full description on the dataset page: https://huggingface.co/datasets/humairmunirawn/UrduG2P.
