CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sk-community /romanized_hindi Romanized Hindi Dataset Dataset Description The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation. It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods. The dataset is designed for training and evaluating Hindi↔Roman transliteration models. Language(s): Hindi, Romanized Hindi Size: ~1.82M rows License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.text1M<n<10M0 likes106 downloads1y agoHugging Face02aneesh-b /SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset. It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api. tabular1K<n<10K0 likes103 downloads4y agoHugging Face03Paul /hatecheck-hindi Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.tabulartext-classification1K<n<10K1 likes99 downloads4y agoHugging Face04Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes85 downloads1y agoHugging Face05ganeshjcs /hindi-article-summarization Summary hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.texttext-generation10K<n<100K0 likes57 downloads3y agoHugging Face06codebyam /Hinglish-Hindi-Transliteration-Datasetgated Hinglish-Hindi Transliteration Dataset We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.texttext-generation1K<n<10K2 likes43 downloads1y agoHugging Face07Process-Venue /Movie_Review_Sentiment_Hinditexttext-classification1K<n<10K0 likes37 downloads10mo agoHugging Face08sepidmnorozy /Hindi_sentimenttextn<1K1 likes32 downloads4y agoHugging Face09Huzayfah-Patel /mindbridge-phq9-hindi-seeds MindBridge Hindi PHQ-9/GAD-7 — Gold Seeds (144 rows) Hand-authored Hindi seeds for PHQ-9 + GAD-7 screening across three personas (postnatal_mother, older_woman, man) in 1:1:1 distribution. Authored via SuperWhisper Scribe with cloud LLM post-process; all rows human-reviewed with review_status=accepted. This seed set drives Phase B teacher expansion (in-context exemplars for Gemma 4 26B-A4B MoE on Vertex MaaS) plus 24 Item-9 (suicidality) extras authored separately. See companion… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-seeds.tabulartext-classificationn<1K0 likes32 downloads1mo agoHugging Face10InfoBayAI /Hindi-Non-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains. The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K1 likes31 downloads9d agoHugging Face11harshitkaran /Hinditext100K<n<1M3 likes30 downloads4y agoHugging Face12Speech-data /Hindi-Speech-Dataset 🎧 Hindi Speech Dataset The Hindi Speech Dataset is a high-quality and structured speech audio dataset developed to support modern AI systems that rely on diverse audio data and scalable voice data. It contains 132 hours of recordings distributed across 565 files, available in MP3 and WAV formats, with a total size of 101 MB. This carefully curated audio dataset provides balanced speaker representation with 49% female and 51% male contributors, covering an age range from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hindi-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes30 downloads6mo agoHugging Face13ud-nlp /hindi-speech-recognition-dataset Hindi Telephone Dialogues Dataset - 760 Hours Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes29 downloads8mo agoHugging Face14miraiminds /function-calling-dataset-Hindi-englishtext1K<n<10K0 likes28 downloads2y agoHugging Face15ganeshjcs /hindi-headline-article-generation Summary hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.texttext-generation100K<n<1M1 likes27 downloads3y agoHugging Face16akashuee /hindi_visual_genometabular10K<n<100K0 likes27 downloads1y agoHugging Face17UniDataPro /hindi-speech-recognition-dataset Hindi Speech Dataset for recognition task Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes26 downloads1mo agoHugging Face18HydraIndicLM /Hindi_Train_ClosedDomainQAThe dataset is the Hindi-only and processed version of https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.hi https://huggingface.co/datasets/xtreme https://huggingface.co/datasets/xquad https://huggingface.co/datasets/databricks/databricks-dolly-15k/viewer/default/train?p=17&f[category][value]=%27closed_qa%27 (closed-qa only) textquestion-answering10K<n<100K0 likes25 downloads3y agoHugging Face19ReySajju742 /Hindi-Poetry-Dataset Hindi Transliteration of Urdu Poetry Dataset Welcome to the Hindi Transliteration of Urdu Poetry Dataset! This dataset features Hindi transliterations of traditional Urdu poetry. Each entry in the dataset includes two columns: Title: The transliterated title of the poem in Hindi. Poem: The transliterated text of the Urdu poem rendered in Hindi script. This dataset is perfect for researchers and developers working on cross-script language processing, transliteration models, and… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Hindi-Poetry-Dataset.texttext-classification1K<n<10K0 likes25 downloads2y agoHugging Face20CodeWithSomesh /english-hindi-vocab-flashcardsgatedtexttext-classification1K<n<10K2 likes25 downloads1y agoHugging Face21d0r1h /HindiNewSummaries How to use this dataset # You can load data using following code and then split into train and validation set from datasets import load_dataset data = load_dataset("d0r1h/HindiNewSummaries") Licensing Note: This license applies to the dataset curation and summaries. Original article copyrights remain with their publishers. This dataset contains news articles and summaries collected from publicly available news websites. The original copyright of the article text… See the full description on the dataset page: https://huggingface.co/datasets/d0r1h/HindiNewSummaries.textsummarization100K<n<1M0 likes24 downloads4mo agoHugging Face22VishalMysore /Hindi_Mithai Dataset Card for Indian Sweets textn<1K0 likes23 downloads3y agoHugging Face23mlexplorer008 /hin_dialect_classificationtext1K<n<10K0 likes23 downloads2y agoHugging Face24bajpaideeksha /english-hindi-colloquial-datasetA curated dataset of colloquial English phrases and their corresponding Hindi translations. This dataset focuses on informal language, including slang, idioms, and everyday expressions, making it ideal for training models that handle casual conversations. Dataset Details: Size:e.g., 500+ phrase pairs] Source: Collected from publicly available conversational datasets, social media, and crowdsourced contributions. Language Pair: English → Hindi Annotations: Each phrase pair is manually verified… See the full description on the dataset page: https://huggingface.co/datasets/bajpaideeksha/english-hindi-colloquial-dataset.texttranslationn<1K2 likes22 downloads2y agoHugging Face25iam-tsr /hindi-sentimentslabel: {'Neutral': 0, 'Positive': 1, 'Negative': 2} texttext-classification100K<n<1M0 likes21 downloads7mo agoHugging Face26shivam9980 /inshorts-hinditext10K<n<100K0 likes19 downloads3y agoHugging Face27TokenBender /sentence_retrieval_hindi_SFTtext10K<n<100K2 likes18 downloads3y agoHugging Face28guneetsk99 /hindi_instruction_set_187Ktext100K<n<1M2 likes18 downloads3y agoHugging Face29wong132 /bengali-hindi-number-blindspot Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion Summary This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step. IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.texttext-generationn<1K0 likes17 downloads7mo agoHugging Face30utkarsharora100 /google_go_emotions_hindi_translatedtabular100K<n<1M1 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.