datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.sphragis
Sphragis
Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc)
prose and verse authorship attribution (AA). Its input is the complete curated
union of the human-annotated CoNLL-U
trees in the supported treebank projects, published in two syntax layers: the
merged human annotation (conllu_human) and one uniform machine parse of every
sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution
tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.sphragis-metre
Sphragis Metre
Sphragis Metre is the scanned-line companion to
Urdatorn/sphragis. It
supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line
units combining human metrical annotation with uniform automatic dependency
annotation. Every curated Hypotactic passage is parsed with the pinned
Ericu950/Stoicheia-tagger-parser
checkpoint.
Tasks
There are three task sizes, 1, 5 and 10 lines, on each of three tracks.
Track
What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.so101-leader-urdf
SO-101 leader URDF
Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration.
Changes
Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions.
Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.ouhd-l-online-urdu-nastaliq-handwriting
OUHD-L: Online Urdu Nastaliq Handwriting — Line Pen Trajectories
Unmodified mirror. This repository re-hosts the OUHD-L v1.0 core release
exactly as published on Zenodo, byte-for-byte. Nothing has been added to or
removed from the data. It exists only to provide an alternative download
endpoint. The canonical source and citation is the Zenodo record:
https://zenodo.org/records/20642162 — DOI
10.5281/zenodo.20642162, version 1.0.0.
Overview
2,403 handwritten Urdu… See the full description on the dataset page: https://huggingface.co/datasets/saad2002/ouhd-l-online-urdu-nastaliq-handwriting.so101_pick_cube_test_20260823_40ep_urdfThis dataset was created using LeRobot.
What this is
The URDF-aligned repair of SN0429/so101_pick_cube_test_20260823_40ep,
for 3D visualisation and Isaac Sim.
LeRobot's degree zero means "the pose held during calibration"; the SO-101 URDF's zero is a different
configuration. On elbow_flex the two are about 97 degrees apart and opposite in sign, so a
recording in LeRobot's convention renders in the wrong posture. This copy maps the angles into the
URDF's frame.
Joint… See the full description on the dataset page: https://huggingface.co/datasets/SN0429/so101_pick_cube_test_20260823_40ep_urdf.UrduSpeech-IndicVoices-ST-kProcessedns-urdu-dataseturdu_finepdfs
What’s inside
data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards.
scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text).
README.md — this file.
If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.URD
license: apache-2.0
Dataset Overview
“Data is not about volume; it is about density.”
This dataset was synthesized using PROD-V2, a high-performance data refinery built to maximize quality density rather than raw volume.
The system treats dataset construction as a multi-objective optimization problem, balancing:
Semantic Entropy (diversity)
Reward Alignment Score (quality)
Noise is removed using geometric filtering, semantic stratification, and discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SofiTesfay2010/URD.urdu-eng-merged-dataset---
language:
- ur
- en
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
tags:
- urdu
- english
- speech
- audio
- asr
- tts
size_categories:
- 10K<n<100K
---
# Urdu + English merged speech dataset
A merged dataset with:
- **Urdu rows**: duration filter + normalization + cleaning
- **English rows**: duration filter only, no text normalization
## Dataset Summary
| Field | Value |
|---|---|
| **Samples** | 241,149 |
| **Total audio** | 488.9 hours |
|… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-eng-merged-dataset.LORU-ED_Roman-Urdu-Event-Detection
LORU-ED: Labeled Original Roman Urdu Dataset for Event Detection
Dataset Description
LORU-ED is a novel Roman Urdu dataset for event detection, containing 7,694
annotated sentences balanced equally between event and non-event classes.
Dataset Structure
Split
Size
Train
6,155
Test
1,539
Features
id: Unique identifier
text: Roman Urdu sentence
language: Language tag (roman_urdu/mixed/en)
is_event: Binary label (1=event, 0=non-event)… See the full description on the dataset page: https://huggingface.co/datasets/Adeel1percentdotcom/LORU-ED_Roman-Urdu-Event-Detection.Urdu-Munch-Processedurdu-dacvae---
language:
- ur
- en
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
tags:
- urdu
- speech
- audio
- asr
- tts
size_categories:
- 10K<n<100K
---
# urdu-dacvae
A cleaned, normalised Urdu (+ limited English) speech dataset derived from
multiple sources, intended for ASR / TTS / VAE latent modelling.
## Dataset Summary
| Field | Value |
|---|---|
| **Samples** | 182,472 |
| **Total audio** | 445.1 hours |
| **Duration filter** | 2.0s < duration < 20.0s… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-dacvae.metanova_testUrdu-News
[Your Dataset Name]
Dataset Description
This dataset appears to be a collection of news headlines and their corresponding news text. Based on the provided image sample, the text content is in a language that uses the Arabic/Persian script, likely Persian (Farsi) or a similar Middle Eastern language. The dataset is structured in a tabular format suitable for various natural language processing tasks related to news content.
Dataset Structure
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-News.Urdu-PDUrdu-Munch-Curriculum-hard
Urdu Munch Curriculum Hard
Long-form and complex utterances for robust TTS training.
Split
train: 269981 rows
Curriculum Rule
audio_token_len >= 223
Notes
Filtered from zuhri025/Urdu-Munch-Processed-s-merged
for long-form / hard TTS training.
urdfstudio:)
Urdu_munch-MyLinafrom datasets import load_dataset
from linacodec.codec import LinaCodec
from IPython.display import Audio
import torch
from datasets import load_dataset
ds = load_dataset("zuhri025/Urdu_munch-MyLina", split="train")
print(ds)
print(ds.column_names)
Pick a sample
sample = ds[0]
Device
device = "cuda" if torch.cuda.is_available() else "cpu"
Convert to tensors and move to device
speech_tokens = torch.tensor(sample["speech_tokens"]).to(device)
global_embedding =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/Urdu_munch-MyLina.Urdu-Munch-Curriculum-easy
Urdu Munch Curriculum Easy
Filtered Urdu TTS training subset for easier, shorter audio sequences.
Split
train: 242549 rows
Curriculum Rule
audio_token_len < 120
Fields
id
transcript
voice
text
timestamp
duration
audio_content_token_indices
audio_global_embedding
audio_token_len
Notes
This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning.
Urdu-Munch-Curriculum-next
Urdu Munch Curriculum Next
Filtered Urdu TTS training subset for the next curriculum stage.
Split
train: 552470 rows
Curriculum Rule
120 <= audio_token_len < 223
Fields
id
transcript
voice
text
timestamp
duration
audio_content_token_indices
audio_global_embedding
audio_token_len
Notes
This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning.
common-voice-urdu-descriptionsUrdu-Munch-Processedurdu_fineweb-2UrduSpeech-IndicVoices-kProcessed-Cleanedcommon-voice-urdu-processed-tagscommon-voice-urdu-tts-text-tagsurdu-asr-multitask-dataset
Urdu Multi-Task ASR Dataset
Preprocessed Urdu dataset for multi-task learning: ASR + Emotion + Gender
Dataset Details
Language: Urdu (ur)
Tasks: ASR (all samples), Emotion (subset), Gender (subset)
Audio: Raw 16kHz mono WAV (ready for wav2vec2 fine-tuning)
Total Samples: 7,155
Splits
Split
Samples
Emotion %
Gender %
Train
5,724
41.9%
48.8%
Validation
715
41.8%
48.8%
Test
716
41.9%
48.9%
Features
audio: Raw audio array (16kHz… See the full description on the dataset page: https://huggingface.co/datasets/abidanoaman/urdu-asr-multitask-dataset.urduAya_dataset
