CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mavkif /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Mavkif/Roman-Urdu-Parl-split.texttranslation1M<n<10M5 likes218 downloads2y agoHugging Face02abdullaharoon /Urdu-Multi-Domain-Benchmark Urdu Multi-Domain Datasets 33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time. Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore. Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts. Permanent archive: Zenodo DOI 10.5281/zenodo.22195610. How to load Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.texttext-classification100K<n<1M0 likes130 downloads24d agoHugging Face03saad2002 /ouhd-l-online-urdu-nastaliq-handwriting OUHD-L: Online Urdu Nastaliq Handwriting — Line Pen Trajectories Unmodified mirror. This repository re-hosts the OUHD-L v1.0 core release exactly as published on Zenodo, byte-for-byte. Nothing has been added to or removed from the data. It exists only to provide an alternative download endpoint. The canonical source and citation is the Zenodo record: https://zenodo.org/records/20642162 — DOI 10.5281/zenodo.20642162, version 1.0.0. Overview 2,403 handwritten Urdu… See the full description on the dataset page: https://huggingface.co/datasets/saad2002/ouhd-l-online-urdu-nastaliq-handwriting.tabular1K<n<10K0 likes109 downloads3d agoHugging Face04abeeranajam31 /urdu-spam-dataset Urdu Spam Detection Dataset Description This dataset is designed for classifying Urdu text into: 0 → Not Spam 1 → Spam It is intended for AI-powered emergency helpline systems (e.g., 1122/911) to filter prank or irrelevant calls. Dataset Structure Format: CSV Column Type Description text string Urdu sentence label int (0/1) Spam classification Example text,label آپ کو میں نے پہلے بھی کال کیا تھا کیا یاد ہے,1 یہ… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/urdu-spam-dataset.texttext-classification1K<n<10K1 likes94 downloads27d agoHugging Face05NeerjaK /Urdu_DataThis repo has cleaned urdu data scraped from the web. text100K<n<1M0 likes88 downloads2y agoHugging Face06Ehtisham1328 /urdu-idioms-with-english-translationtexttranslation1K<n<10K5 likes77 downloads3y agoHugging Face07fatymahaly /urdu_rag_dataset.csv Dataset Card for Urdu RAG Knowledge Base Dataset Overview This dataset is designed specifically to bootstrap and evaluate Retrieval-Augmented Generation (RAG) applications, search systems, and semantic retrieval pipelines using the Urdu language. It contains 185 clean, structured, and informative text chunks covering a wide array of domains. Language: Urdu (ur) Script: Nastaliq / Arabic script (Unicode UTF-8) Total Rows: 185 chunks Format: CSV (id, title… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/urdu_rag_dataset.csv.textn<1K1 likes76 downloads15d agoHugging Face08ReySajju742 /Urdu-Poetry-Dataset Urdu Poetry Dataset Welcome to the Urdu Poetry Dataset! This dataset is a collection of Urdu poems where each entry includes two columns: Title: The title of the poem. Poem: The full text of the poem. This dataset is ideal for natural language processing tasks such as text generation, language modeling, or cultural studies in computational linguistics. Dataset Overview The Urdu Poetry Dataset comprises a diverse collection of Urdu poems, capturing the rich heritage… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-Poetry-Dataset.texttoken-classification1K<n<10K1 likes69 downloads2y agoHugging Face09fatymahaly /Roman-Urdu-Sentiment-Dataset Roman Urdu Sentiment Dataset This repository contains a curated dataset of Roman Urdu text collected from social media interactions, comments, and daily online conversations. Each text entry is paired with a sentiment label for Natural Language Processing (NLP) tasks such as sentiment analysis. Dataset Structure The dataset is formatted in comma-separated values (.csv) with the following columns: text: The sentence, phrase, or social media comment written in… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Roman-Urdu-Sentiment-Dataset.textn<1K0 likes58 downloads16d agoHugging Face10fatymahaly /Urdu_Sentement Roman Urdu Sentiment Analysis Dataset A clean, structured dataset for sentiment analysis in Roman Urdu, split into training and testing sets. This dataset is optimized for quick integration with the Hugging Face datasets library and is ideal for fine-tuning text classification models (such as DistilBERT, mBERT, or XLM-RoBERTa) to handle Roman Urdu customer feedback, reviews, and social media text. Dataset Structure The dataset contains short user reviews and text… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Urdu_Sentement.textn<1K0 likes54 downloads16d agoHugging Face11Khubaib01 /roman-urdu-sentiment-embeddings Roman Urdu Sentiment Embeddings Dataset Overview This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation. Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.text10K<n<100K2 likes40 downloads9mo agoHugging Face12Qasim522 /Roman-Urdu-Parl-split Roman Urdu Parallel Dataset - Split This dataset is just another version of Roman-Urdu-Parl dataset split into train, validation and test set properly. Details follow below. This repository contains a split version of the Roman-Urdu Parallel Dataset (Roman-Urdu-Parl) structured specifically to facilitate fair evaluation in machine transliteration tasks between Urdu and Roman-Urdu. The Roman-Urdu language lacks standard orthography, leading to a wide range of transliteration… See the full description on the dataset page: https://huggingface.co/datasets/Qasim522/Roman-Urdu-Parl-split.texttranslation1M<n<10M0 likes39 downloads10mo agoHugging Face13Madu786 /multilingual-urdu-romanurdu-arabic-english-sentiment Multilingual Sentiment Classification Dataset (EN, UR, Roman UR, AR) Overview This dataset is a clean and balanced multilingual sentiment classification dataset covering four languages: English Urdu Roman Urdu Arabic The dataset is designed to support sentiment analysis and text classification tasks, especially for low-resource languages such as Urdu and Roman Urdu. Sentiment Classes Each text sample belongs to one of the following sentiment categories:… See the full description on the dataset page: https://huggingface.co/datasets/Madu786/multilingual-urdu-romanurdu-arabic-english-sentiment.text100K<n<1M0 likes30 downloads9mo agoHugging Face14Urwashanza /Poly-FEVER-Urdu-Translation Poly-FEVER Urdu Translation Dataset Description This dataset is an Urdu language extension of the original Poly-FEVER benchmark — a multilingual fact verification dataset for hallucination detection in Large Language Models. Urdu (اردو) is spoken by over 230 million people worldwide but was not included in the original Poly-FEVER dataset, which covers 11 languages. This dataset fills that gap by providing complete translations of all 77,971 factual claims… See the full description on the dataset page: https://huggingface.co/datasets/Urwashanza/Poly-FEVER-Urdu-Translation.texttext-classification10K<n<100K0 likes29 downloads3mo agoHugging Face15PuristanLabs1 /Urdu-Turn-Detection-10k Urdu Turn Detection Dataset 🗣️ A high-quality dataset of 10,000 Urdu sentences labeled for Turn Detection (End-of-Turn). This dataset is designed to help conversational AI systems determine if a user has finished speaking (Complete) or is pausing/trailing off (Incomplete). Dataset Details Total Samples: 10,000 Language: Urdu (ur) - Nastaliq/Arabic Script only. Cleanliness: - 100% Urdu Script (No Roman/English). Avg. Sentence Length: - ~7.7 words (33 characters)… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/Urdu-Turn-Detection-10k.texttext-classification10K<n<100K0 likes26 downloads10mo agoHugging Face16Maisum-Abbas-123 /Urdu-Multimodal-Emotion-Datasetaudio1K<n<10K0 likes25 downloads8mo agoHugging Face17hassan4830 /urdu-binary-classification-dataThis Urdu sentiment dataset was formed by concatenating the following two datasets: https://github.com/MuhammadYaseenKhan/Urdu-Sentiment-Corpus https://www.kaggle.com/datasets/akkefa/imdb-dataset-of-50k-movie-translated-urdu-reviews text10K<n<100K1 likes24 downloads4y agoHugging Face18Aimlab /Sentiment-Analysis-Roman-Urdutext10K<n<100K6 likes22 downloads4y agoHugging Face19Mudasir692 /english-urdutext10K<n<100K0 likes21 downloads2y agoHugging Face20atahirsh /urdu-wordstext1K<n<10K1 likes21 downloads1mo agoHugging Face21ReySajju742 /English-Urdu-Dataset English Transliteration (Roman Urdu) of Urdu Poetry Dataset Welcome to the English Transliteration (Roman Urdu) of Urdu Poetry Dataset! This dataset provides a collection of Urdu poetry transliterated into Roman script, making the rich literary heritage of Urdu accessible to a broader audience. Each entry in the dataset includes two columns: Title: The transliterated title of the poem in Roman Urdu. Poem: The full text of the poem in Roman Urdu. This dataset is ideal for… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/English-Urdu-Dataset.texttext-classification1K<n<10K0 likes20 downloads2y agoHugging Face22umar178 /UrduMultiDomainClassification Urdu Multi-Domain Text Classification Dataset Dataset Summary This dataset is a multi-domain Urdu text classification dataset designed for sentiment analysis, intent recognition, topic classification, and binary relevance detection.It contains short Urdu sentences covering multiple real-world domains such as health, education, population, and general/other topics. Each example is annotated with four labels: Sentiment → positive, negative, neutral Topic → health… See the full description on the dataset page: https://huggingface.co/datasets/umar178/UrduMultiDomainClassification.text10K<n<100K0 likes17 downloads1y agoHugging Face23DGurgurov /urdu_sa Sentiment Analysis Data for the Urdu Language Dataset Description: This dataset contains a sentiment analysis dataset from Khan et al. (2020). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{khan2017harnessing, title={Harnessing English Sentiment Lexicons for Polarity Detection in Urdu Tweets: A Baseline Approach}, author={Khan, Muhammad Yaseen and Emaduddin, Shah Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/urdu_sa.texttext-classification10K<n<100K0 likes16 downloads2y agoHugging Face24Omarrran /Sentence_wise_urdu_text_dataset Sentence_wise_urdu_text_dataset Dataset Overview File Information Size: 5.29 MB (5,545,229 bytes) Encoding: UTF-8 Basic Statistics Total Characters: 3,136,348 Total Characters (excluding spaces): 2,472,408 Total Lines: 69,743 Total Words: 666,907 Linguistic Analysis Vocabulary Size: 29,888 Average Word Length: 3.56 characters Median Word Length: 3 characters Average Paragraph Length: 670091.00 words Hapax Legomena… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/Sentence_wise_urdu_text_dataset.texttext-classification10K<n<100K1 likes16 downloads2y agoHugging Face25urdof7 /metanova_testtabular10M<n<100M0 likes16 downloads2y agoHugging Face26ReySajju742 /Urdu-News [Your Dataset Name] Dataset Description This dataset appears to be a collection of news headlines and their corresponding news text. Based on the provided image sample, the text content is in a language that uses the Arabic/Persian script, likely Persian (Farsi) or a similar Middle Eastern language. The dataset is structured in a tabular format suitable for various natural language processing tasks related to news content. Dataset Structure The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Urdu-News.tabularquestion-answering100K<n<1M0 likes16 downloads1y agoHugging Face27Harsit /xnli2.0_train_urdulanguage: ["Urdu"] text100K<n<1M0 likes13 downloads4y agoHugging Face28DANI001 /urdu_ds2300texttext-classification1K<n<10K0 likes12 downloads2y agoHugging Face29amtellezfernandez /urdfstudio:) tabularn<1K0 likes12 downloads11mo agoHugging Face30humairmunirawn /UrduG2P Zuhri — Urdu G2P Dataset Zuhri is a comprehensive and manually verified Urdu Grapheme-to-Phoneme (G2P) dataset. It is designed to aid research and development in areas such as speech synthesis, pronunciation modeling, and computational linguistics, specifically for the Urdu language. This dataset provides accurate phoneme transcriptions and IPA representations, making it ideal for use in building high-quality TTS (Text-to-Speech), ASR (Automatic Speech Recognition), and other… See the full description on the dataset page: https://huggingface.co/datasets/humairmunirawn/UrduG2P.texttext-to-audio10K<n<100K0 likes12 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.