datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.sphragis
Sphragis
Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc)
prose and verse authorship attribution (AA). Its input is the complete curated
union of the human-annotated CoNLL-U
trees in the supported treebank projects, published in two syntax layers: the
merged human annotation (conllu_human) and one uniform machine parse of every
sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution
tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.Vast-Urdu
Vast Urdu Parallel Corpus
Dataset Description
Vast-Urdu is a large-scale collection of parallel text corpora specifically filtered to support Urdu (UR) language research. This dataset was extracted from the liboaccn/nmt-parallel-corpus to provide a dedicated resource for Neural Machine Translation (NMT), cross-lingual understanding, and token-classification tasks involving Urdu.
Source Data
The data is sourced from a massive web-scale crawl, containing… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Vast-Urdu.sphragis-metre
Sphragis Metre
Sphragis Metre is the scanned-line companion to
Urdatorn/sphragis. It
supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line
units combining human metrical annotation with uniform automatic dependency
annotation. Every curated Hypotactic passage is parsed with the pinned
Ericu950/Stoicheia-tagger-parser
checkpoint.
Tasks
There are three task sizes, 1, 5 and 10 lines, on each of three tracks.
Track
What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.sangraha-urdu-onlyurdu_asr_dataroman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.sangraha-urdu-LATN-SYNurdu-ocr-1M
Urdu OCR Dataset (1.5 Million Samples)
Dataset Summary
This is a large-scale synthetic dataset for Urdu Optical Character Recognition (OCR), featuring a groundbreaking Nastaliq collection and a robust Naskh base.
Nastaliq (Primary): 499,845 samples rendered with authentic Jameel Noori Nastaliq ligatures using a custom Chromium-based rendering pipeline.
Naskh: 1,000,160 samples in standard Urdu fonts for baseline OCR tasks.
Totaling 1.5 Million samples, this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-ocr-1M.Indo-Aryan-hin-urd-guj-pan_trainUrduShers-10kRoman-Urdu-Parl-split
Roman Urdu Parallel Dataset - Split
This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below.
This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu.
The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.so101-leader-urdf
SO-101 leader URDF
Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration.
Changes
Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions.
Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.roman_urdu
Dataset Card for Roman Urdu Dataset
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Urdu
Dataset Structure
[More Information Needed]
Data Instances
Wah je wah,Positive,
Data Fields
Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.UrduMMLU
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
Ahmer Tabassum*1 ·
Sarfraz Ahmad*1 ·
Hasan Iqbal*1 ·
Owais Aijaz1 ·
Momina Ahsan1 ·
Preslav Nakov1
1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · *Equal contribution
UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice
questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.munch_urdu_preview
🎧 Munch Preview Dataset
📖 Table of Contents
Dataset Description
Dataset Structure
Dataset Creation
Usage
Considerations
CitationContact
📋 Dataset Description
Overview
Munch Preview is a carefully curated preview dataset containing ** high-quality Urdu text-to-speech samples** from both versions of the Munch dataset family. This lightweight version allows researchers, developers, and practitioners to quickly explore and prototype with… See the full description on the dataset page: https://huggingface.co/datasets/humair025/munch_urdu_preview.Urdu-Finetuning-Data-VibeVoice-Largeurdu_fake_news
Dataset Card for Bend the Truth (Urdu Fake News)
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
news: a string in urdu
label: the label indicating whethere the provided news is real or fake.
category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.imdb_urdu_reviews
Dataset Card for ImDB Urdu Reviews
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
sentence: The movie review which was translated into Urdu.
sentiment: The sentiment exhibited in the review, either positive or negative.
Data Splits
[More… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/imdb_urdu_reviews.AncientGreek-no-sphragis
AncientGreek-no-sphragis (second derivative)
A contamination-controlled derivative of
Ericu950/AncientGreek
at revision 6ac90787c669a7e9218d6d4675a029fa3f10ed99, for pretraining models
that are evaluated on Urdatorn/sphragis
(revision 1e6d8b58d956e84aec7c1c778bef036bd0286fa9) and
Urdatorn/sphragis-metre
(revision 43af5a8b230af1c62d2cafb69c0e3d4d83b81800). Both quality
tiers are retained.
The first derivative removed only source lines that equalled a benchmark unit
after… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/AncientGreek-no-sphragis.urdu-translated-coco-captions-subset
Research Paper: https://www.arxiv.org/abs/2509.09014
Github: https://github.com/umair-hassan2/COCO-Urdu
Overview
Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. COCO-Urdu addresses this gap by providing 59K images and 319K high-quality Urdu captions. Captions were generated via zero-shot translation using SeamlessM4T v2, validated with a hybrid QE pipeline combining COMET-Kiwi, CLIP-based visual grounding, and BERTScore… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset.task1035_pib_translation_tamil_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1035_pib_translation_tamil_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1035_pib_translation_tamil_urdu.pair_hindi_urdu_ipaqaari-0.1-ocr-urdu-news-dataset-smalloga-conllu-stoicheia
OGA parsed with Stoicheia
Period-delimited sentences from the source=oga records of the pristine split of
Ericu950/AncientGreek. A literal . terminates a sentence and is retained. All source columns
are inherited by each sentence. author contains the canonical author label corresponding to the
TLG/CTS code in id; text contains the sentence and conllu contains the Universal
Dependencies-style output from Ericu950/Stoicheia-tagger-parser.
The dataset contains 1,135,428 sentences… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/oga-conllu-stoicheia.urd_ara_Arab_traincommon-voice-urdu-processed-expanded
🎙️ Common Voice Urdu (Processed & Expanded)
The largest ready-to-use Urdu speech dataset for fine-tuning ASR models
Mozilla Common Voice → Preprocessed & Expanded → Whisper-Ready ✨
📊 Dataset at a Glance
Split
Samples
Use
🏋️ Train
54,891
Model training
🔧 Validation
5,000
Hyperparameter tuning
🧪 Test
5,000
Final evaluation
Total
64,891
💡 Audio is pre-resampled to 16kHz — plug directly into Whisper!
📈 3.7x more training data than the… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed-expanded.urdu-audiodataset
Dataset Card for AudioDataset-15
Dataset Description
Dataset Summary
The dataset in question is an audio dataset consisting of recordings in the Urdu language. It has been sourced from Mozilla's Common Voice, a publicly available voice dataset that relies on the contributions of volunteers from various parts of the world. The primary purpose of this dataset is to support the development of voice applications by providing a valuable resource for training machine… See the full description on the dataset page: https://huggingface.co/datasets/HowMannyMore/urdu-audiodataset.urdu-tts-corpus
Urdu TTS Corpus
This dataset is a curated collection of Urdu speech-text pairs, designed for training Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models. It consolidates multiple high-quality sources into a standardized format.
Source Attrbution
This corpus is a merger of the following datasets:
gondal_urdu_tts: muhammadsaadgondal/urdu-tts
urdu_tts_16k: codewithdark/urdu-tts-16000Hz
mozilla_cv_urdu_24: Mozilla Foundation
urdu_tts_fast:… See the full description on the dataset page: https://huggingface.co/datasets/ahmedjaved812/urdu-tts-corpus.pan_Guru_urd_Arab_train
