CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cetiennec /so101-leader-urdf SO-101 leader URDF Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration. Changes Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions. Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.3dn<1K0 likes243 downloads7d agoHugging Face02keplersystems /UrduShers-10ktext10K<n<100K1 likes226 downloads2y agoHugging Face03MBZUAI /UrduMMLU UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding Ahmer Tabassum*1 &nbsp;·&nbsp; Sarfraz Ahmad*1 &nbsp;·&nbsp; Hasan Iqbal*1 &nbsp;·&nbsp; Owais Aijaz1 &nbsp;·&nbsp; Momina Ahsan1 &nbsp;·&nbsp; Preslav Nakov1 1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) &nbsp;·&nbsp; *Equal contribution UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.textquestion-answering10K<n<100K4 likes182 downloads4mo agoHugging Face04Urdatorn /norma Norma Syllabarum Graecarum - A Benchmark for grc Syllabification and Vowel Length Annotation We introduce Norma as a common benchmark for the evaluation and comparison of NLP tools concerning markup of two tasks for Ancient Greek (grc): (1) vowel length of dichronic vowels (alpha, iota, ypsilon) in open syllables (where they impact syllable weight) and (2) syllabification, both boundaries and weight. This means that the benchmark also indirectly tests handling of sandhi… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/norma.text1K<n<10K0 likes60 downloads2mo agoHugging Face05sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes55 downloads9d agoHugging Face06abdullah693 /adaption-urdu-edu-cultural-reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-urdu_edu_cultural_reasoning This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face07abeeranajam31 /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-calls.texttext-generation1K<n<10K0 likes51 downloads28d agoHugging Face08nassimjp /pashto-urdutext1K<n<10K0 likes46 downloads1mo agoHugging Face09Ashar086 /roman-urdu-qwen25-3b-blindspot Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct) Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu). Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4. Evaluation condition pass n rate english 7 8 0.88 formal_urdu 3 8 0.38 roman_urdu 1 8 0.12 Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json Roman Urdu traces: 01_ro: NADRA described as a motor-vehicle department 02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.texttext-generationn<1K0 likes42 downloads2d agoHugging Face10Redgerd /roman-urdu-alpaca-qa-mix Dataset Card for Roman Urdu + Alpaca QA Mix This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total: 500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API. 500 examples in English randomly sampled from the Stanford Alpaca dataset. The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.textquestion-answering1K<n<10K0 likes36 downloads1y agoHugging Face11abeeranajam31 /urdu-emergency-corpus Urdu Emergency Communication Corpus (UEC) A small, balanced, annotated pilot corpus of simulated Urdu emergency-communication utterances, built to study which linguistic features distinguish low- from high-urgency communication using corpus-linguistic methods (frequency, keyness, collocation analysis). Full project, code, executed analysis notebooks, and research report: github.com/abeeranajam31/urdu-emergency-corpus ⚠️ Important: this is a simulated… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-corpus.texttext-classificationn<1K0 likes35 downloads15d agoHugging Face12keplersystems /UrduGhazals-25ktext10K<n<100K0 likes33 downloads2y agoHugging Face13Khurram123 /urdu-poetry-mega-corpus 📜 Urdu Poetry Mega Corpus This dataset is a comprehensive collection of classical and modern Urdu poetry, meticulously curated for training Large Language Models (LLMs) like Qwen, Llama, and Mistral to generate authentic Urdu Ghazals and Nazms. 🌟 Dataset Overview The Urdu Poetry Mega Corpus contains over 42,639 high-quality Urdu couplets (ash'aar). It is designed to capture the structural nuances, rhythmic patterns (Beher), and stylistic essence of renowned poets such… See the full description on the dataset page: https://huggingface.co/datasets/Khurram123/urdu-poetry-mega-corpus.text10K<n<100K0 likes32 downloads7mo agoHugging Face14mira-iitjmu /ns-urdu-datasetaudio1K<n<10K0 likes29 downloads5mo agoHugging Face15Tensoic /GPTeacher-Urdutext10K<n<100K0 likes28 downloads3y agoHugging Face16hamza-amin /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety incidents The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.texttext-generation1K<n<10K1 likes25 downloads9mo agoHugging Face17nassimjp /Liquid-Urdu-Reasoning-Chat-Dataset Liquid Urdu Reasoning Chat Dataset This dataset contains high-quality, synthetically generated Urdu conversational and reasoning data designed for Supervised Fine-Tuning (SFT) of Small Language Models (SLMs), specifically optimized for reasoning models like those run via llama.cpp using DeepSeek-style reasoning formats. Dataset Structure The dataset is formatted in JSONL where each line contains chat history including system prompts, user queries, model… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Liquid-Urdu-Reasoning-Chat-Dataset.text1K<n<10K0 likes25 downloads2d agoHugging Face18SofiTesfay2010 /URD license: apache-2.0 Dataset Overview “Data is not about volume; it is about density.” This dataset was synthesized using PROD-V2, a high-performance data refinery built to maximize quality density rather than raw volume. The system treats dataset construction as a multi-objective optimization problem, balancing: Semantic Entropy (diversity) Reward Alignment Score (quality) Noise is removed using geometric filtering, semantic stratification, and discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SofiTesfay2010/URD.tabular10K<n<100K0 likes23 downloads10mo agoHugging Face19rmahesh /EAMCET-Urdu-Examstextn<1K0 likes22 downloads2y agoHugging Face20fahdmirzac /urdu_bollywood_songs_dataset Bollywood-Inspired Dataset: Movies and Songs Created by Fahd Mirza = https://www.youtube.com/@fahdmirza Overview This dataset is a creative collection of fictional Bollywood movie titles paired with equally fictional song lyrics. Inspired by the rich tradition of Bollywood cinema, where music plays a pivotal role in storytelling, this dataset aims to provide a unique resource for exploring the interplay between movie themes and their musical expressions. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fahdmirzac/urdu_bollywood_songs_dataset.textn<1K0 likes17 downloads3y agoHugging Face21mariasaif20 /Urdu_MultiHop_QAtext10K<n<100K0 likes16 downloads1mo agoHugging Face22keplersystems /UrduPoetry-35ktext10K<n<100K0 likes13 downloads2y agoHugging Face23Almanships /Urdu-Training-for-NLP Urdu Instruction Dataset for NLP A manually curated dataset of 578 Urdu instruction-response pairs for fine-tuning language models on Urdu NLP tasks. Dataset Description This dataset was created to address the lack of instruction-tuning data for Urdu, a low-resource language spoken by over 230 million people. All examples were written and verified by a native Urdu speaker. Dataset Structure Each example contains a conversation with a user… See the full description on the dataset page: https://huggingface.co/datasets/Almanships/Urdu-Training-for-NLP.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face24Safwanahmad619 /adaption-urdu-agri-qa This dataset is a remastered version prepared using Adaption's Adaptive Data platform. urdu_agri_qa This dataset contains question-answer pairs in Urdu focused on agricultural practices, crop diseases, and farming techniques specific to Pakistan. The content covers topics such as wheat rust identification, fertilizer application, garlic cultivation, and organic farming opportunities. Each entry provides concise, actionable advice for farmers regarding plant health and yield… See the full description on the dataset page: https://huggingface.co/datasets/Safwanahmad619/adaption-urdu-agri-qa.textn<1K0 likes11 downloads5mo agoHugging Face25zuhri025 /Urdu_munch-MyLinafrom datasets import load_dataset from linacodec.codec import LinaCodec from IPython.display import Audio import torch from datasets import load_dataset ds = load_dataset("zuhri025/Urdu_munch-MyLina", split="train") print(ds) print(ds.column_names) Pick a sample sample = ds[0] Device device = "cuda" if torch.cuda.is_available() else "cpu" Convert to tensors and move to device speech_tokens = torch.tensor(sample["speech_tokens"]).to(device) global_embedding =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/Urdu_munch-MyLina.tabular100K<n<1M0 likes10 downloads9mo agoHugging Face26hassan7272 /urdu-finance-qa 🇵🇰 Urdu Financial QA Dataset (Roman Urdu + Urdu + Mixed) 🚀 First open-source Urdu financial QA dataset focused on Pakistan + Islamic finance 🚀 A high-quality, domain-specific Urdu financial dataset for Pakistan, combining Urdu script, Roman Urdu, and code-mixed queries, designed for real-world NLP systems and RAG applications. 📌 Overview This dataset contains 1,510 carefully curated question-answer pairs focused on financial scenarios relevant to Pakistani users… See the full description on the dataset page: https://huggingface.co/datasets/hassan7272/urdu-finance-qa.textquestion-answering1K<n<10K0 likes10 downloads5mo agoHugging Face27Maitreyajayaraj /hf_medical_debug_api_v5_urdu_full.jsontextn<1K0 likes9 downloads5mo agoHugging Face28Maitreyajayaraj /hf_medical_debug_compiler_v5_urdu_batch2.jsontextn<1K0 likes9 downloads5mo agoHugging Face29Maitreyajayaraj /hf_medical_debug_critical_api_v7_urdu_elite.jsontextn<1K0 likes9 downloads5mo agoHugging Face30Safwanahmad619 /adaption-urdu-agri-qa-v1 This dataset is a remastered version prepared using Adaption's Adaptive Data platform. urdu_agri_qa This dataset contains question-answer pairs in Urdu focused on agricultural practices, crop diseases, and farming techniques specific to Pakistan. The content covers topics such as wheat rust identification, fertilizer application, garlic cultivation, and organic farming opportunities. Each entry provides concise, actionable advice for farmers regarding plant health and yield… See the full description on the dataset page: https://huggingface.co/datasets/Safwanahmad619/adaption-urdu-agri-qa-v1.textn<1K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.