CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face02Urdatorn /sphragis Sphragis Sphragis (σφραγίς, "sigil") is a benchmark for Ancient Greek (grc) prose and verse authorship attribution (AA). Its input is the complete curated union of the human-annotated CoNLL-U trees in the supported treebank projects, published in two syntax layers: the merged human annotation (conllu_human) and one uniform machine parse of every sentence (conllu_machine). It defines 1-, 5-, and 10-sentence attribution tasks on six tracks. The complementary scanned-line… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis.tabulartext-classification100K<n<1M0 likes933 downloads2d agoHugging Face03Urdatorn /sphragis-metre Sphragis Metre Sphragis Metre is the scanned-line companion to Urdatorn/sphragis. It supports Ancient Greek authorship attribution from exact 1-, 5-, and 10-line units combining human metrical annotation with uniform automatic dependency annotation. Every curated Hypotactic passage is parsed with the pinned Ericu950/Stoicheia-tagger-parser checkpoint. Tasks There are three task sizes, 1, 5 and 10 lines, on each of three tracks. Track What its rows are… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-metre.tabulartext-classification100K<n<1M0 likes385 downloads2d agoHugging Face04cetiennec /so101-leader-urdf SO-101 leader URDF Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration. Changes Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions. Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.3dn<1K0 likes210 downloads7d agoHugging Face05saad2002 /ouhd-l-online-urdu-nastaliq-handwriting OUHD-L: Online Urdu Nastaliq Handwriting — Line Pen Trajectories Unmodified mirror. This repository re-hosts the OUHD-L v1.0 core release exactly as published on Zenodo, byte-for-byte. Nothing has been added to or removed from the data. It exists only to provide an alternative download endpoint. The canonical source and citation is the Zenodo record: https://zenodo.org/records/20642162 — DOI 10.5281/zenodo.20642162, version 1.0.0. Overview 2,403 handwritten Urdu… See the full description on the dataset page: https://huggingface.co/datasets/saad2002/ouhd-l-online-urdu-nastaliq-handwriting.tabular1K<n<10K0 likes109 downloads3d agoHugging Face06SN0429 /so101_pick_cube_test_20260823_40ep_urdfThis dataset was created using LeRobot. What this is The URDF-aligned repair of SN0429/so101_pick_cube_test_20260823_40ep, for 3D visualisation and Isaac Sim. LeRobot's degree zero means "the pose held during calibration"; the SO-101 URDF's zero is a different configuration. On elbow_flex the two are about 97 degrees apart and opposite in sign, so a recording in LeRobot's convention renders in the wrong posture. This copy maps the angles into the URDF's frame. Joint… See the full description on the dataset page: https://huggingface.co/datasets/SN0429/so101_pick_cube_test_20260823_40ep_urdf.tabularrobotics10K<n<100K0 likes79 downloads26d agoHugging Face07humair025 /UrduSpeech-IndicVoices-ST-kProcessedtabular10K<n<100K0 likes31 downloads7mo agoHugging Face08mira-iitjmu /ns-urdu-datasetaudio1K<n<10K0 likes29 downloads5mo agoHugging Face09humair025 /urdu_finepdfs What’s inside data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards. scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text). README.md — this file. If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.tabulartext-classification100K<n<1M0 likes26 downloads10mo agoHugging Face10SofiTesfay2010 /URD license: apache-2.0 Dataset Overview “Data is not about volume; it is about density.” This dataset was synthesized using PROD-V2, a high-performance data refinery built to maximize quality density rather than raw volume. The system treats dataset construction as a multi-objective optimization problem, balancing: Semantic Entropy (diversity) Reward Alignment Score (quality) Noise is removed using geometric filtering, semantic stratification, and discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SofiTesfay2010/URD.tabular10K<n<100K0 likes22 downloads10mo agoHugging Face11zuhri025 /urdu-eng-merged-dataset--- language: - ur - en license: cc-by-4.0 task_categories: - automatic-speech-recognition tags: - urdu - english - speech - audio - asr - tts size_categories: - 10K<n<100K --- # Urdu + English merged speech dataset A merged dataset with: - **Urdu rows**: duration filter + normalization + cleaning - **English rows**: duration filter only, no text normalization ## Dataset Summary | Field | Value | |---|---| | **Samples** | 241,149 | | **Total audio** | 488.9 hours | |… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-eng-merged-dataset.tabular100K<n<1M0 likes20 downloads4mo agoHugging Face12Adeel1percentdotcom /LORU-ED_Roman-Urdu-Event-Detection LORU-ED: Labeled Original Roman Urdu Dataset for Event Detection Dataset Description LORU-ED is a novel Roman Urdu dataset for event detection, containing 7,694 annotated sentences balanced equally between event and non-event classes. Dataset Structure Split Size Train 6,155 Test 1,539 Features id: Unique identifier text: Roman Urdu sentence language: Language tag (roman_urdu/mixed/en) is_event: Binary label (1=event, 0=non-event)… See the full description on the dataset page: https://huggingface.co/datasets/Adeel1percentdotcom/LORU-ED_Roman-Urdu-Event-Detection.tabulartext-classification1K<n<10K0 likes19 downloads4mo agoHugging Face13zuhri025 /Urdu-Munch-Processedtabular10K<n<100K0 likes18 downloads8mo agoHugging Face14zuhri025 /urdu-dacvae--- language: - ur - en license: cc-by-4.0 task_categories: - automatic-speech-recognition tags: - urdu - speech - audio - asr - tts size_categories: - 10K<n<100K --- # urdu-dacvae A cleaned, normalised Urdu (+ limited English) speech dataset derived from multiple sources, intended for ASR / TTS / VAE latent modelling. ## Dataset Summary | Field | Value | |---|---| | **Samples** | 182,472 | | **Total audio** | 445.1 hours | | **Duration filter** | 2.0s < duration < 20.0s… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-dacvae.tabular100K<n<1M0 likes18 downloads4mo agoHugging Face15urdof7 /metanova_testtabular10M<n<100M0 likes16 downloads2y agoHugging Face16ReySajju742 /Urdu-News [Your Dataset Name] Dataset Description This dataset appears to be a collection of news headlines and their corresponding news text. Based on the provided image sample, the text content is in a language that uses the Arabic/Persian script, likely Persian (Farsi) or a similar Middle Eastern language. The dataset is structured in a tabular format suitable for various natural language processing tasks related to news content. Dataset Structure The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-News.tabularquestion-answering100K<n<1M0 likes16 downloads1y agoHugging Face17PleIAs /Urdu-PDtabular1K<n<10K0 likes14 downloads2y agoHugging Face18zuhri025 /Urdu-Munch-Curriculum-hard Urdu Munch Curriculum Hard Long-form and complex utterances for robust TTS training. Split train: 269981 rows Curriculum Rule audio_token_len >= 223 Notes Filtered from zuhri025/Urdu-Munch-Processed-s-merged for long-form / hard TTS training. tabulartext-to-speech100K<n<1M0 likes13 downloads5mo agoHugging Face19amtellezfernandez /urdfstudio:) tabularn<1K0 likes12 downloads11mo agoHugging Face20zuhri025 /Urdu_munch-MyLinafrom datasets import load_dataset from linacodec.codec import LinaCodec from IPython.display import Audio import torch from datasets import load_dataset ds = load_dataset("zuhri025/Urdu_munch-MyLina", split="train") print(ds) print(ds.column_names) Pick a sample sample = ds[0] Device device = "cuda" if torch.cuda.is_available() else "cpu" Convert to tensors and move to device speech_tokens = torch.tensor(sample["speech_tokens"]).to(device) global_embedding =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/Urdu_munch-MyLina.tabular100K<n<1M0 likes12 downloads9mo agoHugging Face21zuhri025 /Urdu-Munch-Curriculum-easy Urdu Munch Curriculum Easy Filtered Urdu TTS training subset for easier, shorter audio sequences. Split train: 242549 rows Curriculum Rule audio_token_len < 120 Fields id transcript voice text timestamp duration audio_content_token_indices audio_global_embedding audio_token_len Notes This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning. tabulartext-to-speech100K<n<1M0 likes12 downloads5mo agoHugging Face22zuhri025 /Urdu-Munch-Curriculum-next Urdu Munch Curriculum Next Filtered Urdu TTS training subset for the next curriculum stage. Split train: 552470 rows Curriculum Rule 120 <= audio_token_len < 223 Fields id transcript voice text timestamp duration audio_content_token_indices audio_global_embedding audio_token_len Notes This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning. tabulartext-to-speech100K<n<1M0 likes11 downloads5mo agoHugging Face23saleem0044 /common-voice-urdu-descriptionstabular10K<n<100K0 likes10 downloads1y agoHugging Face24humair025 /Urdu-Munch-Processedtabular10K<n<100K0 likes10 downloads8mo agoHugging Face25humair025 /urdu_fineweb-2tabular1M<n<10M1 likes9 downloads1y agoHugging Face26humair025 /UrduSpeech-IndicVoices-kProcessed-Cleanedtabular100K<n<1M0 likes9 downloads7mo agoHugging Face27saleem0044 /common-voice-urdu-processed-tagstabular10K<n<100K0 likes8 downloads1y agoHugging Face28saleem0044 /common-voice-urdu-tts-text-tagstabular10K<n<100K0 likes7 downloads1y agoHugging Face29abidanoaman /urdu-asr-multitask-dataset Urdu Multi-Task ASR Dataset Preprocessed Urdu dataset for multi-task learning: ASR + Emotion + Gender Dataset Details Language: Urdu (ur) Tasks: ASR (all samples), Emotion (subset), Gender (subset) Audio: Raw 16kHz mono WAV (ready for wav2vec2 fine-tuning) Total Samples: 7,155 Splits Split Samples Emotion % Gender % Train 5,724 41.9% 48.8% Validation 715 41.8% 48.8% Test 716 41.9% 48.9% Features audio: Raw audio array (16kHz… See the full description on the dataset page: https://huggingface.co/datasets/abidanoaman/urdu-asr-multitask-dataset.tabular1K<n<10K0 likes7 downloads8mo agoHugging Face30PuristanLabs /urduAya_datasettabularn<1K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.