CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASLP-lab /UrduSpeech Dataset Summary UrduSpeech is a large-scale, high-fidelity Urdu speech corpus comprising 156 hours of audio with comprehensive 12-dimensional paralinguistic metadata. The corpus addresses the critical under-resourcing of Urdu in speech technology by providing: 71,792 diarized utterances across diverse content categories Three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2h), Urdu-English Code-Switched (US-CS, 89.4h), and Pakistani-Accented English (US-EngPk, 7.3h)… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/UrduSpeech.audio10K<n<100K7 likes2.9k downloads4mo agoHugging Face02Behavision /URDF 3D Model URDF Dataset This is a URDF dataset of 3D models, with both textured and untextured versions, designed to support research in robotics simulation, grasping, and physics simulation. Dataset Description This dataset consists of two parts, totaling 500 models: 235 Textured URDF Models: This part includes detailed texture maps, suitable for scenarios requiring high-fidelity rendering. 265 Untextured URDF Models: This part focuses on the physical and geometric… See the full description on the dataset page: https://huggingface.co/datasets/Behavision/URDF.3dn<1K3 likes1.9k downloads1y agoHugging Face03haixuantao /urdfs0 likes1.4k downloads6mo agoHugging Face04humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face05Urdatorn /sphragis Sphragis Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc) prose and verse authorship attribution (AA). Its input is the complete curated union of the human-annotated CoNLL-U trees in the supported treebank projects, published in two syntax layers: the merged human annotation (conllu_human) and one uniform machine parse of every sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.tabulartext-classification100K<n<1M0 likes933 downloads2d agoHugging Face06CSALT /deepfake_detection_dataset_urdu Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset". The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.audio1K<n<10K5 likes755 downloads2y agoHugging Face07LocalWorldModels /partnet-mobility-urdf-parts0 likes478 downloads1y agoHugging Face08ReySajju742 /Vast-Urdu Vast Urdu Parallel Corpus Dataset Description Vast-Urdu is a large-scale collection of parallel text corpora specifically filtered to support Urdu (UR) language research. This dataset was extracted from the liboaccn/nmt-parallel-corpus to provide a dedicated resource for Neural Machine Translation (NMT), cross-lingual understanding, and token-classification tasks involving Urdu. Source Data The data is sourced from a massive web-scale crawl, containing… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Vast-Urdu.texttranslation10M<n<100M0 likes412 downloads8mo agoHugging Face09Urdatorn /sphragis-metre Sphragis Metre Sphragis Metre is the scanned-line companion to Urdatorn/sphragis. It supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line units combining human metrical annotation with uniform automatic dependency annotation. Every curated Hypotactic passage is parsed with the pinned Ericu950/Stoicheia-tagger-parser checkpoint. Tasks There are three task sizes, 1, 5 and 10 lines, on each of three tracks. Track What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.tabulartext-classification100K<n<1M0 likes385 downloads2d agoHugging Face10humairmunirawn /sangraha-urdu-onlytext1M<n<10M0 likes376 downloads5mo agoHugging Face11adnankarim /urdu_asr_dataaudio10K<n<100K2 likes332 downloads3y agoHugging Face12community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes328 downloads2y agoHugging Face13azeem-ahmed /Common_Voice_Corpus_22_0_Urdu Common Voice Corpus 22.0 - Urdu This dataset contains the Urdu subset of the Mozilla Common Voice 22.0 corpus, released in June 2025.It consists of crowdsourced speech recordings and their corresponding text transcriptions, collected to support open-source speech technology. Dataset Summary The Common Voice Corpus 22.0 Urdu dataset provides high-quality speech data for automatic speech recognition (ASR), speaker identification, and linguistic research in Urdu.It includes… See the full description on the dataset page: https://huggingface.co/datasets/azeem-ahmed/Common_Voice_Corpus_22_0_Urdu.audioautomatic-speech-recognition1 likes320 downloads1y agoHugging Face14humairmunirawn /sangraha-urdu-LATN-SYNtext1M<n<10M0 likes277 downloads5mo agoHugging Face15PuristanLabs1 /urdu-ocr-1M Urdu OCR Dataset (1.5 Million Samples) Dataset Summary This is a large-scale synthetic dataset for Urdu Optical Character Recognition (OCR), featuring a groundbreaking Nastaliq collection and a robust Naskh base. Nastaliq (Primary): 499,845 samples rendered with authentic Jameel Noori Nastaliq ligatures using a custom Chromium-based rendering pipeline. Naskh: 1,000,160 samples in standard Urdu fonts for baseline OCR tasks. Totaling 1.5 Million samples, this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-ocr-1M.imageimage-to-text1M<n<10M6 likes276 downloads8mo agoHugging Face16yiyic /Indo-Aryan-hin-urd-guj-pan_traintext1M<n<10M0 likes259 downloads2y agoHugging Face17keplersystems /UrduShers-10ktext10K<n<100K1 likes228 downloads2y agoHugging Face18Idrees0 /Urdu-Speech-AI-Audiosaudio1K<n<10K0 likes223 downloads2mo agoHugging Face19Mavkif /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.texttranslation1M<n<10M5 likes218 downloads2y agoHugging Face20cetiennec /so101-leader-urdf SO-101 leader URDF Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration. Changes Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions. Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.3dn<1K0 likes210 downloads7d agoHugging Face21community-datasets /urdu_sentiment_corpus“Urdu Sentiment Corpus” (USC) shares the dat of Urdu tweets for the sentiment analysis and polarity detection. The dataset is consisting of tweets and overall, the dataset is comprising over 17, 185 tokens with 52% records as positive, and 48 % records as negative.text-classification1K<n<10K1 likes203 downloads3y agoHugging Face22kaab4321 /UrduTTS UrduTTS A Studio-Quality Urdu Speech Corpus with Urdu, Phonemized, and Romanized Transcriptions 91.9 hours · 57,873 utterances · 44.1 kHz · 3 aligned text representations Dataset Summary UrduTTS is the largest openly available Urdu text-to-speech corpus with three aligned text representations for every utterance: native Urdu script, Phonemized (IPA) text, and Romanized (Latin) text. Urdu is spoken by roughly 250 million people but is badly… See the full description on the dataset page: https://huggingface.co/datasets/kaab4321/UrduTTS.text-to-speech10K<n<100K0 likes194 downloads14d agoHugging Face23community-datasets /roman_urdu Dataset Card for Roman Urdu Dataset Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages Urdu Dataset Structure [More Information Needed] Data Instances Wah je wah,Positive, Data Fields Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.texttext-classification10K<n<100K3 likes191 downloads2y agoHugging Face24MBZUAI /UrduMMLU UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding Ahmer Tabassum*1 &nbsp;·&nbsp; Sarfraz Ahmad*1 &nbsp;·&nbsp; Hasan Iqbal*1 &nbsp;·&nbsp; Owais Aijaz1 &nbsp;·&nbsp; Momina Ahsan1 &nbsp;·&nbsp; Preslav Nakov1 1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) &nbsp;·&nbsp; *Equal contribution UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.textquestion-answering10K<n<100K4 likes186 downloads4mo agoHugging Face25kingabzpro /Urdu-ASR-flags Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/kingabzpro/Urdu-ASR-flags.0 likes185 downloads11mo agoHugging Face26humair025 /munch_urdu_preview 🎧 Munch Preview Dataset 📖 Table of Contents Dataset Description Dataset Structure Dataset Creation Usage Considerations CitationContact 📋 Dataset Description Overview Munch Preview is a carefully curated preview dataset containing ** high-quality Urdu text-to-speech samples** from both versions of the Munch dataset family. This lightweight version allows researchers, developers, and practitioners to quickly explore and prototype with… See the full description on the dataset page: https://huggingface.co/datasets/humair025/munch_urdu_preview.audiotext-to-speech1K<n<10K1 likes179 downloads10mo agoHugging Face27Maisum-Abbas-123 /Urdu-Finetuning-Data-VibeVoice-Largeaudio10K<n<100K0 likes179 downloads8mo agoHugging Face28Shaahadath /Noorwise-quran-urdu-max Noorwise Quran Urdu Max Complete Quran (114 Surahs) with Urdu and Hindi verse-by-verse translation audio. Each surah contains Arabic recitation followed by Urdu and Hindi translation. Dataset Summary Total Surahs: 114 (complete Quran) Language: Arabic (recitation) + Urdu + Hindi (translation) Audio Format: MP3 Artist: The Quran DVD Total Duration: ~42 hours Total Size: ~2.1 GB License: MIT Data Structure Each entry in Urdumax.json contains the… See the full description on the dataset page: https://huggingface.co/datasets/Shaahadath/Noorwise-quran-urdu-max.audion<1K0 likes175 downloads1mo agoHugging Face29community-datasets /urdu_fake_news Dataset Card for Bend the Truth (Urdu Fake News) Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields news: a string in urdu label: the label indicating whethere the provided news is real or fake. category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.texttext-classificationn<1K3 likes168 downloads2y agoHugging Face30mirfan899 /imdb_urdu_reviews Dataset Card for ImDB Urdu Reviews Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields sentence: The movie review which was translated into Urdu. sentiment: The sentiment exhibited in the review, either positive or negative. Data Splits [More… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/imdb_urdu_reviews.texttext-classification10K<n<100K1 likes163 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.