datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UrduSpeech
Dataset Summary
UrduSpeech is a large-scale, high-fidelity Urdu speech corpus comprising 156 hours of audio with comprehensive 12-dimensional paralinguistic metadata. The corpus addresses the critical under-resourcing of Urdu in speech technology by providing:
71,792 diarized utterances across diverse content categories
Three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2h), Urdu-English Code-Switched (US-CS, 89.4h), and Pakistani-Accented English (US-EngPk, 7.3h)… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/UrduSpeech.URDF
3D Model URDF Dataset
This is a URDF dataset of 3D models, with both textured and untextured versions, designed to support research in robotics simulation, grasping, and physics simulation.
Dataset Description
This dataset consists of two parts, totaling 500 models:
235 Textured URDF Models: This part includes detailed texture maps, suitable for scenarios requiring high-fidelity rendering.
265 Untextured URDF Models: This part focuses on the physical and geometric… See the full description on the dataset page: https://huggingface.co/datasets/Behavision/URDF.urdfsUrdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.sphragis
Sphragis
Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc)
prose and verse authorship attribution (AA). Its input is the complete curated
union of the human-annotated CoNLL-U
trees in the supported treebank projects, published in two syntax layers: the
merged human annotation (conllu_human) and one uniform machine parse of every
sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution
tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.deepfake_detection_dataset_urdu
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset
This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset".
The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.partnet-mobility-urdf-partsVast-Urdu
Vast Urdu Parallel Corpus
Dataset Description
Vast-Urdu is a large-scale collection of parallel text corpora specifically filtered to support Urdu (UR) language research. This dataset was extracted from the liboaccn/nmt-parallel-corpus to provide a dedicated resource for Neural Machine Translation (NMT), cross-lingual understanding, and token-classification tasks involving Urdu.
Source Data
The data is sourced from a massive web-scale crawl, containing… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Vast-Urdu.sphragis-metre
Sphragis Metre
Sphragis Metre is the scanned-line companion to
Urdatorn/sphragis. It
supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line
units combining human metrical annotation with uniform automatic dependency
annotation. Every curated Hypotactic passage is parsed with the pinned
Ericu950/Stoicheia-tagger-parser
checkpoint.
Tasks
There are three task sizes, 1, 5 and 10 lines, on each of three tracks.
Track
What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.sangraha-urdu-onlyurdu_asr_dataroman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.Common_Voice_Corpus_22_0_Urdu
Common Voice Corpus 22.0 - Urdu
This dataset contains the Urdu subset of the Mozilla Common Voice 22.0 corpus, released in June 2025.It consists of crowdsourced speech recordings and their corresponding text transcriptions, collected to support open-source speech technology.
Dataset Summary
The Common Voice Corpus 22.0 Urdu dataset provides high-quality speech data for automatic speech recognition (ASR), speaker identification, and linguistic research in Urdu.It includes… See the full description on the dataset page: https://huggingface.co/datasets/azeem-ahmed/Common_Voice_Corpus_22_0_Urdu.sangraha-urdu-LATN-SYNurdu-ocr-1M
Urdu OCR Dataset (1.5 Million Samples)
Dataset Summary
This is a large-scale synthetic dataset for Urdu Optical Character Recognition (OCR), featuring a groundbreaking Nastaliq collection and a robust Naskh base.
Nastaliq (Primary): 499,845 samples rendered with authentic Jameel Noori Nastaliq ligatures using a custom Chromium-based rendering pipeline.
Naskh: 1,000,160 samples in standard Urdu fonts for baseline OCR tasks.
Totaling 1.5 Million samples, this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-ocr-1M.Indo-Aryan-hin-urd-guj-pan_trainUrduShers-10kUrdu-Speech-AI-AudiosRoman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.so101-leader-urdf
SO-101 leader URDF
Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration.
Changes
Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions.
Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.urdu_sentiment_corpus“Urdu Sentiment Corpus” (USC) shares the dat of Urdu tweets for the sentiment analysis and polarity detection.
The dataset is consisting of tweets and overall, the dataset is comprising over 17, 185 tokens
with 52% records as positive, and 48 % records as negative.UrduTTS
UrduTTS
A Studio-Quality Urdu Speech Corpus with Urdu, Phonemized, and Romanized Transcriptions
91.9 hours · 57,873 utterances · 44.1 kHz · 3 aligned text representations
Dataset Summary
UrduTTS is the largest openly available Urdu text-to-speech corpus with three aligned text
representations for every utterance: native Urdu script, Phonemized (IPA) text, and
Romanized (Latin) text.
Urdu is spoken by roughly 250 million people but is badly… See the full description on the dataset page: https://huggingface.co/datasets/kaab4321/UrduTTS.roman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.UrduMMLU
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
Ahmer Tabassum*1 ·
Sarfraz Ahmad*1 ·
Hasan Iqbal*1 ·
Owais Aijaz1 ·
Momina Ahsan1 ·
Preslav Nakov1
1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · *Equal contribution
UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice
questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.Urdu-ASR-flags
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/Urdu-ASR-flags.munch_urdu_preview
🎧 Munch Preview Dataset
📖 Table of Contents
Dataset Description
Dataset Structure
Dataset Creation
Usage
Considerations
CitationContact
📋 Dataset Description
Overview
Munch Preview is a carefully curated preview dataset containing ** high-quality Urdu text-to-speech samples** from both versions of the Munch dataset family. This lightweight version allows researchers, developers, and practitioners to quickly explore and prototype with… See the full description on the dataset page: https://huggingface.co/datasets/humair025/munch_urdu_preview.Urdu-Finetuning-Data-VibeVoice-LargeNoorwise-quran-urdu-max
Noorwise Quran Urdu Max
Complete Quran (114 Surahs) with Urdu and Hindi verse-by-verse translation audio. Each surah contains Arabic recitation followed by Urdu and Hindi translation.
Dataset Summary
Total Surahs: 114 (complete Quran)
Language: Arabic (recitation) + Urdu + Hindi (translation)
Audio Format: MP3
Artist: The Quran DVD
Total Duration: ~42 hours
Total Size: ~2.1 GB
License: MIT
Data Structure
Each entry in Urdumax.json contains the… See the full description on the dataset page: https://huggingface.co/datasets/Shaahadath/Noorwise-quran-urdu-max.urdu_fake_news
Dataset Card for Bend the Truth (Urdu Fake News)
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
news: a string in urdu
label: the label indicating whethere the provided news is real or fake.
category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.imdb_urdu_reviews
Dataset Card for ImDB Urdu Reviews
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
sentence: The movie review which was translated into Urdu.
sentiment: The sentiment exhibited in the review, either positive or negative.
Data Splits
[More… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/imdb_urdu_reviews.
