CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Urdatorn /sphragis-olmo1b-adaptation-corpus Sphragis OLMo-1B adaptation corpus Version-controlled input for adapting allenai/OLMo-1B-hf to Ancient Greek before authorship-language-model training. It contains only OGA whole works whose TLG author occurs in neither Sphragis benchmark. Text has the exact model-facing benchmark surface form: polytonic-aware lowercasing with grc_utils.lower_grc, removal of all editorial punctuation, normalization of whitespace, and removal of consonant-final elision marks. Splits are made over… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/sphragis-olmo1b-adaptation-corpus.texttext-generation1K<n<10K0 likes132 downloads28d agoHugging Face02Lots-of-LoRAs /task1035_pib_translation_tamil_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1035_pib_translation_tamil_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1035_pib_translation_tamil_urdu.texttext-generation1K<n<10K0 likes120 downloads2y agoHugging Face03Lots-of-LoRAs /task990_pib_translation_urdu_marathi Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task990_pib_translation_urdu_marathi Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task990_pib_translation_urdu_marathi.texttext-generation1K<n<10K0 likes110 downloads2y agoHugging Face04Ehtisham1328 /urdu-idioms-with-english-translationtexttranslation1K<n<10K5 likes70 downloads3y agoHugging Face05Lots-of-LoRAs /task1037_pib_translation_telugu_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1037_pib_translation_telugu_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1037_pib_translation_telugu_urdu.texttext-generationn<1K0 likes69 downloads2y agoHugging Face06abdullah693 /adaption-urdu-edu-cultural-reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-urdu_edu_cultural_reasoning This dataset contains a mixed collection of question-answer pairs and linguistic tasks presented in both English and Urdu. The content spans multiple domains including history, biology, geography, and Urdu literature, featuring multiple-choice questions, translation exercises, and poetic composition prompts. Samples include historical treaty analysis… See the full description on the dataset page: https://huggingface.co/datasets/abdullah693/adaption-urdu-edu-cultural-reasoning.texttext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face07abeeranajam31 /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-emergency-calls.texttext-generation1K<n<10K0 likes51 downloads27d agoHugging Face08AhmadMustafa /Urdu-Instruct-News-Article-Generation Dataset Card for "Urdu-Instruct-News-Article-Generation" This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota. Task: Generate the News Article from the given headline. Split Size: train: 100674 test: 11187 Prompt Template (In Urdu): Random.choice b.w these 2. The First template is template_id 1 and the second template is template_id 2 in the dataset. [ "اس دی گی ایک خبر… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Article-Generation.texttext-generation100K<n<1M4 likes45 downloads3y agoHugging Face09Ashar086 /roman-urdu-qwen25-3b-blindspot Roman Urdu / code-switch blind spot (Qwen2.5-3B-Instruct) Hand-built eval: 8 Pakistani situations x 3 surfaces (English, formal Urdu, Roman Urdu). Model: Qwen/Qwen2.5-3B-Instruct. Greedy decoding, 4-bit, Colab T4. Evaluation condition pass n rate english 7 8 0.88 formal_urdu 3 8 0.38 roman_urdu 1 8 0.12 Files: prompts.jsonl, outputs.jsonl, judged.jsonl, scores.json Roman Urdu traces: 01_ro: NADRA described as a motor-vehicle department 02_ro:… See the full description on the dataset page: https://huggingface.co/datasets/Ashar086/roman-urdu-qwen25-3b-blindspot.texttext-generationn<1K0 likes42 downloads2d agoHugging Face10Redgerd /roman-urdu-alpaca-qa-mix Dataset Card for Roman Urdu + Alpaca QA Mix This dataset is intended to support fine-tuning and evaluation of language models that understand and respond to Roman Urdu and English instructions. It consists of 1,022 records in total: 500 examples in Roman Urdu generated from high-quality Urdu sources and transliterated using the ChatGPT API. 500 examples in English randomly sampled from the Stanford Alpaca dataset. The dataset follows the same format as Alpaca-style instruction… See the full description on the dataset page: https://huggingface.co/datasets/Redgerd/roman-urdu-alpaca-qa-mix.textquestion-answering1K<n<10K0 likes36 downloads1y agoHugging Face11AhmadMustafa /Urdu-Instruct-News-Headline-Generation Dataset Card for "Urdu-Instruct-News-Headline-Generation" This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota. Task: Generate the News Headline from the given News. Split Size: train: 100674 test: 11187 Prompt Template (In Urdu): Random.choice b.w these 2. The first template is template_id 1, and the second template is template_id 2 in the dataset. ["اس اردو پیراگراف… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Headline-Generation.texttext-generation100K<n<1M0 likes33 downloads3y agoHugging Face12AhmadMustafa /Urdu-Instruct-News-Category-Classification Dataset Card for "Urdu-Instruct-News-Category-Classification" This Dataset is converted from the original dataset by Khalid Hussain, Nimra Mughal, Irfan Ali, Saif Hassan, Sher Muhammad Daudpota. Task: Generate the News Paragraph, and classify the news category from it. Split Size: train: 100674 test: 11187 Prompt Template (In Urdu): Random.choice b.w these 2. The first template is template_id 1 in the dataset, second template is template_id 2 in… See the full description on the dataset page: https://huggingface.co/datasets/AhmadMustafa/Urdu-Instruct-News-Category-Classification.texttext-classification100K<n<1M0 likes30 downloads3y agoHugging Face13humair025 /urdu_finepdfs What’s inside data/ (optional) — small example files / scripts [FUTURE] . This repo is primarily a pointer + helpers to the official FinePDFs Urdu shards. scripts/ [FUTURE] — utility scripts to list, preview, and filter Urdu parquet shards (example: extract metadata, sample text, convert to plain text). README.md — this file. If you cloned this repo to help with the downstream work, expect the real Urdu shards to be loaded from the official Hugging Face hub (see examples below).… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu_finepdfs.tabulartext-classification100K<n<1M0 likes28 downloads10mo agoHugging Face14hamza-amin /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety incidents The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.texttext-generation1K<n<10K1 likes25 downloads9mo agoHugging Face15Lots-of-LoRAs /task1053_pib_translation_hindi_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1053_pib_translation_hindi_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1053_pib_translation_hindi_urdu.texttext-generation1K<n<10K0 likes24 downloads2y agoHugging Face16muhammadUsman31254 /urdu-english-name-variants Names Dataset (English-Urdu-Variants) This dataset contains names with their standardized English form, Urdu script, and common English variants. Dataset Structure en_std: Standardized English name ur: Name in Urdu script en_var: Common English variants/spellings of the name Usage from datasets import load_dataset dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants") Languages English (primary and variants) Urdu Use… See the full description on the dataset page: https://huggingface.co/datasets/muhammadUsman31254/urdu-english-name-variants.texttext-classification1K<n<10K0 likes22 downloads1y agoHugging Face17Lots-of-LoRAs /task1057_pib_translation_english_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1057_pib_translation_english_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1057_pib_translation_english_urdu.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face18Xhaheen /Alpaca_urdu_2024_1Description The Alpaca Urdu 🦙 is a translation of the original dataset into Urdu. This dataset is a part of the Alpaca project and is designed for NLP tasks. 🌐 Dataset Information Size: The translated dataset contains [45,000] samples. Languages: Urdu License: [cc-by-4.0] Original Dataset: Alpaca Cleaned datasetColumns The translated dataset includes the following columns: input: input text in Urdu. output: translated output in Urdu. answer_lengths: Lengths of the answers. Example Usage… See the full description on the dataset page: https://huggingface.co/datasets/Xhaheen/Alpaca_urdu_2024_1.texttext-generation10K<n<100K1 likes17 downloads3y agoHugging Face19Omarrran /Sentence_wise_urdu_text_dataset Sentence_wise_urdu_text_dataset Dataset Overview File Information Size: 5.29 MB (5,545,229 bytes) Encoding: UTF-8 Basic Statistics Total Characters: 3,136,348 Total Characters (excluding spaces): 2,472,408 Total Lines: 69,743 Total Words: 666,907 Linguistic Analysis Vocabulary Size: 29,888 Average Word Length: 3.56 characters Median Word Length: 3 characters Average Paragraph Length: 670091.00 words Hapax Legomena… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Sentence_wise_urdu_text_dataset.texttext-classification10K<n<100K1 likes16 downloads2y agoHugging Face20PuristanLabs1 /GSM8K_Urdu GSM8K Urdu: Grade School Math Word Problems in Urdu Dataset Description GSM8K Urdu is a high quality Urdu mathematical reasoning dataset replicating GSM8K dataset by OpenAI GSM8K (Grade School Math 8K), containing 6,365 grade school math word problems with step by step reasoning in Urdu script (اردو). This dataset is specifically adapted for Pakistani Islamic cultural context with appropriate content modifications. Dataset Summary Total Examples: 6,365… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/GSM8K_Urdu.textquestion-answering1K<n<10K4 likes16 downloads10mo agoHugging Face21Lots-of-LoRAs /task989_pib_translation_marathi_urdu Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task989_pib_translation_marathi_urdu Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task989_pib_translation_marathi_urdu.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face22Lots-of-LoRAs /task1058_pib_translation_urdu_english Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1058_pib_translation_urdu_english Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1058_pib_translation_urdu_english.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face23feifan961206 /urdu-english-name-variants Names Dataset (English-Urdu-Variants) This dataset contains names with their standardized English form, Urdu script, and common English variants. Dataset Structure en_std: Standardized English name ur: Name in Urdu script en_var: Common English variants/spellings of the name Usage from datasets import load_dataset dataset = load_dataset("muhammadUsman31254/urdu-english-name-variants") Languages English (primary and variants)… See the full description on the dataset page: https://huggingface.co/datasets/feifan961206/urdu-english-name-variants.texttext-classification1K<n<10K0 likes14 downloads3mo agoHugging Face24Almanships /Urdu-Training-for-NLP Urdu Instruction Dataset for NLP A manually curated dataset of 578 Urdu instruction-response pairs for fine-tuning language models on Urdu NLP tasks. Dataset Description This dataset was created to address the lack of instruction-tuning data for Urdu, a low-resource language spoken by over 230 million people. All examples were written and verified by a native Urdu speaker. Dataset Structure Each example contains a conversation with a user… See the full description on the dataset page: https://huggingface.co/datasets/Almanships/Urdu-Training-for-NLP.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face25mahwizzzz /UrduQuotesThe Urdu Quotes Dataset contains a collection of quotes in Urdu. texttext-generation1K<n<10K1 likes12 downloads3y agoHugging Face26HaseebAsif /UrduReason-Eval UrduReason-Eval A Standardized Urdu Reasoning Evaluation Benchmark for Large Language Models UrduReason-Eval is a high-difficulty, evaluation-only reasoning benchmark designed specifically to measure multi-step reasoning capabilities in Urdu-language LLMs. It is one of the few publicly available datasets that jointly evaluates linguistic understanding and formal reasoning in Urdu, a low-resource language spoken by over 230 million people. If you are evaluating Urdu LLM reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HaseebAsif/UrduReason-Eval.textquestion-answeringn<1K1 likes6 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.