datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sphragis-olmo1b-adaptation-corpus
Sphragis OLMo-1B adaptation corpus
Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek
before authorship-language-model training. It contains only OGA whole works
whose TLG author occurs in neither Sphragis benchmark.
Text has the exact model-facing benchmark surface form: polytonic-aware
lowercasing with grc_utils.lower_grc, removal of all editorial punctuation,
normalization of whitespace, and removal of consonant-final elision marks.
Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.task1035_pib_translation_tamil_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1035_pib_translation_tamil_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1035_pib_translation_tamil_urdu.task990_pib_translation_urdu_marathi
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task990_pib_translation_urdu_marathi
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task990_pib_translation_urdu_marathi.urdu-idioms-with-english-translationtask1037_pib_translation_telugu_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1037_pib_translation_telugu_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1037_pib_translation_telugu_urdu.adaption-urdu-edu-cultural-reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-urdu_edu_cultural_reasoning
This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-calls.Urdu-Instruct-News-Article-Generation
Dataset Card for "Urdu-Instruct-News-Article-Generation"
This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota.
Task:
Generate the News Article from the given headline.
Split Size:
train: 100674
test: 11187
Prompt Template (In Urdu):
Random.choice b.w these 2. The First template is template_id 1 and the second template is template_id 2 in the dataset.
[ "اس دی گی ایک خبر… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Article-Generation.roman-urdu-qwen25-3b-blindspot
Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct)
Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu).
Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4.
Evaluation
condition
pass
n
rate
english
7
8
0.88
formal_urdu
3
8
0.38
roman_urdu
1
8
0.12
Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json
Roman Urdu traces:
01_ro: NADRA described as a motor-vehicle department
02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.roman-urdu-alpaca-qa-mix
Dataset Card for Roman Urdu + Alpaca QA Mix
This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total:
500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API.
500 examples in English randomly sampled from the Stanford Alpaca dataset.
The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.Urdu-Instruct-News-Headline-Generation
Dataset Card for "Urdu-Instruct-News-Headline-Generation"
This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota.
Task:
Generate the News Headline from the given News.
Split Size:
train: 100674
test: 11187
Prompt Template (In Urdu):
Random.choice b.w these 2. The first template is template_id 1, and the second template is template_id 2 in the dataset.
["اس اردو پیراگراف… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Headline-Generation.Urdu-Instruct-News-Category-Classification
Dataset Card for "Urdu-Instruct-News-Category-Classification"
This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota.
Task:
Generate the News Paragraph, and classify the news category from it.
Split Size:
train: 100674
test: 11187
Prompt Template (In Urdu):
Random.choice b.w these 2. The first template is template_id 1 in the dataset, second template is template_id 2 in… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Category-Classification.urdu_finepdfs
What’s inside
data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards.
scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text).
README.md — this file.
If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.task1053_pib_translation_hindi_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1053_pib_translation_hindi_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1053_pib_translation_hindi_urdu.urdu-english-name-variants
Names Dataset (English-Urdu-Variants)
This dataset contains names with their standardized English form, Urdu script, and common English variants.
Dataset Structure
en_std: Standardized English name
ur: Name in Urdu script
en_var: Common English variants/spellings of the name
Usage
from datasets import load_dataset
dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants")
Languages
English (primary and variants)
Urdu
Use… See the full description on the dataset page: https://huggingface.co/datasets/muhammadUsman31254/urdu-english-name-variants.task1057_pib_translation_english_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1057_pib_translation_english_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1057_pib_translation_english_urdu.Alpaca_urdu_2024_1Description
The Alpaca Urdu 🦙 is a translation of the original dataset into Urdu. This dataset is a part of the Alpaca project and is designed for NLP tasks. 🌐
Dataset Information
Size: The translated dataset contains [45,000] samples.
Languages: Urdu
License: [cc-by-4.0]
Original Dataset: Alpaca Cleaned datasetColumns
The translated dataset includes the following columns:
input: input text in Urdu.
output: translated output in Urdu.
answer_lengths: Lengths of the answers.
Example Usage… See the full description on the dataset page: https://huggingface.co/datasets/Xhaheen/Alpaca_urdu_2024_1.Sentence_wise_urdu_text_dataset
Sentence_wise_urdu_text_dataset
Dataset Overview
File Information
Size: 5.29 MB (5,545,229 bytes)
Encoding: UTF-8
Basic Statistics
Total Characters: 3,136,348
Total Characters (excluding spaces): 2,472,408
Total Lines: 69,743
Total Words: 666,907
Linguistic Analysis
Vocabulary Size: 29,888
Average Word Length: 3.56 characters
Median Word Length: 3 characters
Average Paragraph Length: 670091.00 words
Hapax Legomena… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Sentence_wise_urdu_text_dataset.GSM8K_Urdu
GSM8K Urdu: Grade School Math Word Problems in Urdu
Dataset Description
GSM8K Urdu is a high quality Urdu mathematical reasoning dataset replicating GSM8K dataset by OpenAI GSM8K (Grade School Math 8K), containing 6,365 grade school math word problems with step by step reasoning in Urdu script (اردو). This dataset is specifically adapted for Pakistani Islamic cultural context with appropriate content modifications.
Dataset Summary
Total Examples: 6,365… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/GSM8K_Urdu.task989_pib_translation_marathi_urdu
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task989_pib_translation_marathi_urdu
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task989_pib_translation_marathi_urdu.task1058_pib_translation_urdu_english
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1058_pib_translation_urdu_english
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1058_pib_translation_urdu_english.urdu-english-name-variants
Names Dataset (English-Urdu-Variants)
This dataset contains names with their standardized English form, Urdu script, and common English variants.
Dataset Structure
en_std: Standardized English name
ur: Name in Urdu script
en_var: Common English variants/spellings of the name
Usage
from datasets import load_dataset
dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants")
Languages
English (primary and variants)… See the full description on the dataset page: https://huggingface.co/datasets/feifan961206/urdu-english-name-variants.Urdu-Training-for-NLP
Urdu Instruction Dataset for NLP
A manually curated dataset of 578 Urdu instruction-response
pairs for fine-tuning language models on Urdu NLP tasks.
Dataset Description
This dataset was created to address the lack of
instruction-tuning data for Urdu, a low-resource language
spoken by over 230 million people. All examples were written
and verified by a native Urdu speaker.
Dataset Structure
Each example contains a conversation with a user… See the full description on the dataset page: https://huggingface.co/datasets/Almanships/Urdu-Training-for-NLP.UrduQuotesThe Urdu Quotes Dataset contains a collection of quotes in Urdu.
UrduReason-Eval
UrduReason-Eval
A Standardized Urdu Reasoning Evaluation Benchmark for Large Language Models
UrduReason-Eval is a high-difficulty, evaluation-only reasoning benchmark designed specifically to measure multi-step reasoning capabilities in Urdu-language LLMs.
It is one of the few publicly available datasets that jointly evaluates linguistic understanding and formal reasoning in Urdu, a low-resource language spoken by over 230 million people.
If you are evaluating Urdu LLM reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HaseebAsif/UrduReason-Eval.
