urd
Datasets
All datasets matching “urd”UrduSpeech
Dataset Summary
UrduSpeech is a large-scale, high-fidelity Urdu speech corpus comprising 156 hours of audio with comprehensive 12-dimensional paralinguistic metadata. The corpus addresses the critical under-resourcing of Urdu in speech technology by providing:
71,792 diarized utterances across diverse content categories
Three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2h), Urdu-English Code-Switched (US-CS, 89.4h), and Pakistani-Accented English (US-EngPk, 7.3h)… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/UrduSpeech.URDF
3D Model URDF Dataset
This is a URDF dataset of 3D models, with both textured and untextured versions, designed to support research in robotics simulation, grasping, and physics simulation.
Dataset Description
This dataset consists of two parts, totaling 500 models:
235 Textured URDF Models: This part includes detailed texture maps, suitable for scenarios requiring high-fidelity rendering.
265 Untextured URDF Models: This part focuses on the physical and geometric… See the full description on the dataset page: https://huggingface.co/datasets/Behavision/URDF.urdfsUrdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.sphragis
Sphragis
Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc)
prose and verse authorship attribution (AA). Its input is the complete curated
union of the human-annotated CoNLL-U
trees in the supported treebank projects, published in two syntax layers: the
merged human annotation (conllu_human) and one uniform machine parse of every
sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution
tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.deepfake_detection_dataset_urdu
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset
This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset".
The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.
