datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so101-leader-urdf
SO-101 leader URDF
Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration.
Changes
Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions.
Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.UrduShers-10kUrduMMLU
UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
Ahmer Tabassum*1 ·
Sarfraz Ahmad*1 ·
Hasan Iqbal*1 ·
Owais Aijaz1 ·
Momina Ahsan1 ·
Preslav Nakov1
1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · *Equal contribution
UrduMMLU is a large-scale, human-curated benchmark of 26,431 multiple-choice
questions written natively in Urdu. Questions are… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/UrduMMLU.norma
Norma Syllabarum Graecarum - A Benchmark for grc Syllabification and Vowel Length Annotation
We introduce Norma as a common benchmark for the evaluation and comparison of NLP tools concerning markup of two tasks for Ancient Greek (grc): (1) vowel length of dichronic vowels (alpha, iota, ypsilon) in open syllables (where they impact syllable weight) and (2) syllabification, both boundaries and weight. This means that the benchmark also indirectly tests handling of sandhi… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/norma.urdu-asr-error-correction-data
Urdu ASR Generative Error Correction Dataset
This dataset contains paired training and testing data for post-ASR error correction in Urdu.
Dataset Details
Language: Urdu (ur)
Task: ASR Error Correction
License: CC BY-NC 4.0
Dataset Structure
The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold).
train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-calls.pashto-urduroman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.roman-urdu-alpaca-qa-mix
Dataset Card for Roman Urdu + Alpaca QA Mix
This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total:
500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API.
500 examples in English randomly sampled from the Stanford Alpaca dataset.
The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.urdu-emergency-corpus
Urdu Emergency Communication Corpus (UEC)
A small, balanced, annotated pilot corpus of simulated Urdu
emergency-communication utterances, built to study which linguistic
features distinguish low- from high-urgency communication using
corpus-linguistic methods (frequency, keyness, collocation analysis).
Full project, code, executed analysis notebooks, and research report:
github.com/abeeranajam31/urdu-emergency-corpus
⚠️ Important: this is a simulated… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-corpus.UrduGhazals-25kurdu-poetry-mega-corpus
📜 Urdu Poetry Mega Corpus
This dataset is a comprehensive collection of classical and modern Urdu poetry, meticulously curated for training Large Language Models (LLMs) like Qwen, Llama, and Mistral to generate authentic Urdu Ghazals and Nazms.
🌟 Dataset Overview
The Urdu Poetry Mega Corpus contains over 42,639 high-quality Urdu couplets (ash'aar). It is designed to capture the structural nuances, rhythmic patterns (Beher), and stylistic essence of renowned poets such… See the full description on the dataset page: https://huggingface.co/datasets/Khurram123/urdu-poetry-mega-corpus.ns-urdu-datasetGPTeacher-Urduurdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.Liquid-Urdu-Reasoning-Chat-Dataset
Liquid Urdu Reasoning Chat Dataset
This dataset contains high-quality, synthetically generated Urdu conversational and reasoning data designed for Supervised Fine-Tuning (SFT) of Small Language Models (SLMs), specifically optimized for reasoning models like those run via llama.cpp using DeepSeek-style reasoning formats.
Dataset Structure
The dataset is formatted in JSONL where each line contains chat history including system prompts, user queries, model… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Liquid-Urdu-Reasoning-Chat-Dataset.URD
license: apache-2.0
Dataset Overview
“Data is not about volume; it is about density.”
This dataset was synthesized using PROD-V2, a high-performance data refinery built to maximize quality density rather than raw volume.
The system treats dataset construction as a multi-objective optimization problem, balancing:
Semantic Entropy (diversity)
Reward Alignment Score (quality)
Noise is removed using geometric filtering, semantic stratification, and discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SofiTesfay2010/URD.EAMCET-Urdu-Examsurdu_bollywood_songs_dataset
Bollywood-Inspired Dataset: Movies and Songs
Created by Fahd Mirza = https://www.youtube.com/@fahdmirza
Overview
This dataset is a creative collection of fictional Bollywood movie titles paired with equally fictional song lyrics. Inspired by the rich tradition of Bollywood cinema, where music plays a pivotal role in storytelling, this dataset aims to provide a unique resource for exploring the interplay between movie themes and their musical expressions.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fahdmirzac/urdu_bollywood_songs_dataset.Urdu_MultiHop_QAUrduPoetry-35kUrdu-Training-for-NLP
Urdu Instruction Dataset for NLP
A manually curated dataset of 578 Urdu instruction-response
pairs for fine-tuning language models on Urdu NLP tasks.
Dataset Description
This dataset was created to address the lack of
instruction-tuning data for Urdu, a low-resource language
spoken by over 230 million people. All examples were written
and verified by a native Urdu speaker.
Dataset Structure
Each example contains a conversation with a user… See the full description on the dataset page: https://huggingface.co/datasets/Almanships/Urdu-Training-for-NLP.adaption-urdu-agri-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
urdu_agri_qa
This dataset contains question-answer pairs in Urdu focused on agricultural practices, crop diseases, and farming techniques specific to Pakistan. The content covers topics such as wheat rust identification, fertilizer application, garlic cultivation, and organic farming opportunities. Each entry provides concise, actionable advice for farmers regarding plant health and yield… See the full description on the dataset page: https://huggingface.co/datasets/Safwanahmad619/adaption-urdu-agri-qa.Urdu_munch-MyLinafrom datasets import load_dataset
from linacodec.codec import LinaCodec
from IPython.display import Audio
import torch
from datasets import load_dataset
ds = load_dataset("zuhri025/Urdu_munch-MyLina", split="train")
print(ds)
print(ds.column_names)
Pick a sample
sample = ds[0]
Device
device = "cuda" if torch.cuda.is_available() else "cpu"
Convert to tensors and move to device
speech_tokens = torch.tensor(sample["speech_tokens"]).to(device)
global_embedding =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/Urdu_munch-MyLina.urdu-finance-qa
🇵🇰 Urdu Financial QA Dataset (Roman Urdu + Urdu + Mixed)
🚀 First open-source Urdu financial QA dataset focused on Pakistan + Islamic finance
🚀 A high-quality, domain-specific Urdu financial dataset for Pakistan, combining Urdu script, Roman Urdu, and code-mixed queries, designed for real-world NLP systems and RAG applications.
📌 Overview
This dataset contains 1,510 carefully curated question-answer pairs focused on financial scenarios relevant to Pakistani users… See the full description on the dataset page: https://huggingface.co/datasets/hassan7272/urdu-finance-qa.hf_medical_debug_api_v5_urdu_full.jsonhf_medical_debug_compiler_v5_urdu_batch2.jsonhf_medical_debug_critical_api_v7_urdu_elite.jsonadaption-urdu-agri-qa-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
urdu_agri_qa
This dataset contains question-answer pairs in Urdu focused on agricultural practices, crop diseases, and farming techniques specific to Pakistan. The content covers topics such as wheat rust identification, fertilizer application, garlic cultivation, and organic farming opportunities. Each entry provides concise, actionable advice for farmers regarding plant health and yield… See the full description on the dataset page: https://huggingface.co/datasets/Safwanahmad619/adaption-urdu-agri-qa-v1.
