CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face02Urdatorn /sphragis Sphragis Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc) prose and verse authorship attribution (AA). Its input is the complete curated union of the human-annotated CoNLL-U trees in the supported treebank projects, published in two syntax layers: the merged human annotation (conllu_human) and one uniform machine parse of every sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.tabulartext-classification100K<n<1M0 likes933 downloads2d agoHugging Face03ReySajju742 /Vast-Urdu Vast Urdu Parallel Corpus Dataset Description Vast-Urdu is a large-scale collection of parallel text corpora specifically filtered to support Urdu (UR) language research. This dataset was extracted from the liboaccn/nmt-parallel-corpus to provide a dedicated resource for Neural Machine Translation (NMT), cross-lingual understanding, and token-classification tasks involving Urdu. Source Data The data is sourced from a massive web-scale crawl, containing… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Vast-Urdu.texttranslation10M<n<100M0 likes412 downloads8mo agoHugging Face04Urdatorn /sphragis-metre Sphragis Metre Sphragis Metre is the scanned-line companion to Urdatorn/sphragis. It supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line units combining human metrical annotation with uniform automatic dependency annotation. Every curated Hypotactic passage is parsed with the pinned Ericu950/Stoicheia-tagger-parser checkpoint. Tasks There are three task sizes, 1, 5 and 10 lines, on each of three tracks. Track What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.tabulartext-classification100K<n<1M0 likes385 downloads2d agoHugging Face05humairmunirawn /sangraha-urdu-onlytext1M<n<10M0 likes376 downloads5mo agoHugging Face06adnankarim /urdu_asr_dataaudio10K<n<100K2 likes332 downloads3y agoHugging Face07community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes328 downloads2y agoHugging Face08humairmunirawn /sangraha-urdu-LATN-SYNtext1M<n<10M0 likes277 downloads5mo agoHugging Face09PuristanLabs1 /urdu-ocr-1M Urdu OCR Dataset (1.5 Million Samples) Dataset Summary This is a large-scale synthetic dataset for Urdu Optical Character Recognition (OCR), featuring a groundbreaking Nastaliq collection and a robust Naskh base. Nastaliq (Primary): 499,845 samples rendered with authentic Jameel Noori Nastaliq ligatures using a custom Chromium-based rendering pipeline. Naskh: 1,000,160 samples in standard Urdu fonts for baseline OCR tasks. Totaling 1.5 Million samples, this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-ocr-1M.imageimage-to-text1M<n<10M6 likes276 downloads8mo agoHugging Face10yiyic /Indo-Aryan-hin-urd-guj-pan_traintext1M<n<10M0 likes259 downloads2y agoHugging Face11keplersystems /UrduShers-10ktext10K<n<100K1 likes228 downloads2y agoHugging Face12Mavkif /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.texttranslation1M<n<10M5 likes218 downloads2y agoHugging Face13cetiennec /so101-leader-urdf SO-101 leader URDF Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration. Changes Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions. Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.3dn<1K0 likes210 downloads7d agoHugging Face14community-datasets /roman_urdu Dataset Card for Roman Urdu Dataset Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages Urdu Dataset Structure [More Information Needed] Data Instances Wah je wah,Positive, Data Fields Each row consists of a short Urdu text, followed by a sentiment label. The labels are one of Positive, Negative, and Neutral. Note that the original source file is a… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu.texttext-classification10K<n<100K3 likes191 downloads2y agoHugging Face15MBZUAI /UrduMMLU UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding Ahmer Tabassum*1 &nbsp;·&nbsp; Sarfraz Ahmad*1 &nbsp;·&nbsp; Hasan Iqbal*1 &nbsp;·&nbsp; Owais Aijaz1 &nbsp;·&nbsp; Momina Ahsan1 &nbsp;·&nbsp; Preslav Nakov1 1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) &nbsp;·&nbsp; *Equal contribution UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.textquestion-answering10K<n<100K4 likes186 downloads4mo agoHugging Face16humair025 /munch_urdu_preview 🎧 Munch Preview Dataset 📖 Table of Contents Dataset Description Dataset Structure Dataset Creation Usage Considerations CitationContact 📋 Dataset Description Overview Munch Preview is a carefully curated preview dataset containing ** high-quality Urdu text-to-speech samples** from both versions of the Munch dataset family. This lightweight version allows researchers, developers, and practitioners to quickly explore and prototype with… See the full description on the dataset page: https://huggingface.co/datasets/humair025/munch_urdu_preview.audiotext-to-speech1K<n<10K1 likes179 downloads10mo agoHugging Face17Maisum-Abbas-123 /Urdu-Finetuning-Data-VibeVoice-Largeaudio10K<n<100K0 likes179 downloads8mo agoHugging Face18community-datasets /urdu_fake_news Dataset Card for Bend the Truth (Urdu Fake News) Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields news: a string in urdu label: the label indicating whethere the provided news is real or fake. category: The intent of the news being presented. The available 5… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/urdu_fake_news.texttext-classificationn<1K3 likes168 downloads2y agoHugging Face19mirfan899 /imdb_urdu_reviews Dataset Card for ImDB Urdu Reviews Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields sentence: The movie review which was translated into Urdu. sentiment: The sentiment exhibited in the review, either positive or negative. Data Splits [More… See the full description on the dataset page: https://huggingface.co/datasets/mirfan899/imdb_urdu_reviews.texttext-classification10K<n<100K1 likes163 downloads2y agoHugging Face20Urdatorn /AncientGreek-no-sphragis AncientGreek-no-sphragis (second derivative) A contamination-controlled derivative of Ericu950/AncientGreek at revision 6ac90787c669a7e9218d6d4675a029fa3f10ed99, for pretraining models that are evaluated on Urdatorn/sphragis (revision 1e6d8b58d956e84aec7c1c778bef036bd0286fa9) and Urdatorn/sphragis-metre (revision 43af5a8b230af1c62d2cafb69c0e3d4d83b81800). Both quality tiers are retained. The first derivative removed only source lines that equalled a benchmark unit after… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/AncientGreek-no-sphragis.textfill-mask1M<n<10M0 likes154 downloads2d agoHugging Face21umairhassan02 /urdu-translated-coco-captions-subset Research Paper: https://www.arxiv.org/abs/2509.09014 Github: https://github.com/umair-hassan2/COCO-Urdu Overview Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. COCO-Urdu addresses this gap by providing 59K images and 319K high-quality Urdu captions. Captions were generated via zero-shot translation using SeamlessM4T v2, validated with a hybrid QE pipeline combining COMET-Kiwi, CLIP-based visual grounding, and BERTScore… See the full description on the dataset page: https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset.image10K<n<100K0 likes153 downloads1y agoHugging Face22Lots-of-LoRAs /task1035_pib_translation_tamil_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1035_pib_translation_tamil_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1035_pib_translation_tamil_urdu.texttext-generation1K<n<10K0 likes146 downloads2y agoHugging Face23iggy12345 /pair_hindi_urdu_ipatext10M<n<100M0 likes144 downloads1y agoHugging Face24oddadmix /qaari-0.1-ocr-urdu-news-dataset-smallimageimage-to-text10K<n<100K1 likes143 downloads5mo agoHugging Face25Urdatorn /oga-conllu-stoicheia OGA parsed with Stoicheia Period-delimited sentences from the source=oga records of the pristine split of Ericu950/AncientGreek. A literal . terminates a sentence and is retained. All source columns are inherited by each sentence. author contains the canonical author label corresponding to the TLG/CTS code in id; text contains the sentence and conllu contains the Universal Dependencies-style output from Ericu950/Stoicheia-tagger-parser. The dataset contains 1,135,428 sentences… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/oga-conllu-stoicheia.texttoken-classification1M<n<10M0 likes139 downloads1mo agoHugging Face26yiyic /urd_ara_Arab_traintext1M<n<10M0 likes137 downloads2y agoHugging Face27khawajaaliarshad /common-voice-urdu-processed-expanded 🎙️ Common Voice Urdu (Processed & Expanded) The largest ready-to-use Urdu speech dataset for fine-tuning ASR models Mozilla Common Voice → Preprocessed & Expanded → Whisper-Ready ✨ 📊 Dataset at a Glance Split Samples Use 🏋️ Train 54,891 Model training 🔧 Validation 5,000 Hyperparameter tuning 🧪 Test 5,000 Final evaluation Total 64,891 💡 Audio is pre-resampled to 16kHz — plug directly into Whisper! 📈 3.7x more training data than the… See the full description on the dataset page: https://huggingface.co/datasets/khawajaaliarshad/common-voice-urdu-processed-expanded.audio10K<n<100K2 likes137 downloads9mo agoHugging Face28HowMannyMore /urdu-audiodataset Dataset Card for AudioDataset-15 Dataset Description Dataset Summary The dataset in question is an audio dataset consisting of recordings in the Urdu language. It has been sourced from Mozilla's Common Voice, a publicly available voice dataset that relies on the contributions of volunteers from various parts of the world. The primary purpose of this dataset is to support the development of voice applications by providing a valuable resource for training machine… See the full description on the dataset page: https://huggingface.co/datasets/HowMannyMore/urdu-audiodataset.audiotranslation10K<n<100K0 likes136 downloads3y agoHugging Face29ahmedjaved812 /urdu-tts-corpus Urdu TTS Corpus This dataset is a curated collection of Urdu speech-text pairs, designed for training Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models. It consolidates multiple high-quality sources into a standardized format. Source Attrbution This corpus is a merger of the following datasets: gondal_urdu_tts: muhammadsaadgondal/urdu-tts urdu_tts_16k: codewithdark/urdu-tts-16000Hz mozilla_cv_urdu_24: Mozilla Foundation urdu_tts_fast:… See the full description on the dataset page: https://huggingface.co/datasets/ahmedjaved812/urdu-tts-corpus.audiotext-to-speech100K<n<1M1 likes136 downloads6mo agoHugging Face30yiyic /pan_Guru_urd_Arab_traintext1M<n<10M0 likes134 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.