CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sk-community /romanized_hindi Romanized Hindi Dataset Dataset Description The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation. It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods. The dataset is designed for training and evaluating Hindi↔Roman transliteration models. Language(s): Hindi, Romanized Hindi Size: ~1.82M rows License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.text1M<n<10M0 likes105 downloads1y agoHugging Face02aneesh-b /SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset. It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api. tabular1K<n<10K0 likes101 downloads4y agoHugging Face03Paul /hatecheck-hindi Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.tabulartext-classification1K<n<10K1 likes94 downloads4y agoHugging Face04Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes86 downloads1y agoHugging Face05ganeshjcs /hindi-article-summarization Summary hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.texttext-generation10K<n<100K0 likes58 downloads3y agoHugging Face06codebyam /Hinglish-Hindi-Transliteration-Datasetgated Hinglish-Hindi Transliteration Dataset We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.texttext-generation1K<n<10K2 likes44 downloads1y agoHugging Face07Process-Venue /Movie_Review_Sentiment_Hinditexttext-classification1K<n<10K0 likes38 downloads10mo agoHugging Face08sepidmnorozy /Hindi_sentimenttextn<1K1 likes33 downloads4y agoHugging Face09Huzayfah-Patel /mindbridge-phq9-hindi-seeds MindBridge Hindi PHQ-9/GAD-7 — Gold Seeds (144 rows) Hand-authored Hindi seeds for PHQ-9 + GAD-7 screening across three personas (postnatal_mother, older_woman, man) in 1:1:1 distribution. Authored via SuperWhisper Scribe with cloud LLM post-process; all rows human-reviewed with review_status=accepted. This seed set drives Phase B teacher expansion (in-context exemplars for Gemma 4 26B-A4B MoE on Vertex MaaS) plus 24 Item-9 (suicidality) extras authored separately. See companion… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-seeds.tabulartext-classificationn<1K0 likes33 downloads29d agoHugging Face10InfoBayAI /Hindi-Non-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains. The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K1 likes33 downloads8d agoHugging Face11harshitkaran /Hinditext100K<n<1M3 likes31 downloads4y agoHugging Face12HydraIndicLM /Hindi_Train_ClosedDomainQAThe dataset is the Hindi-only and processed version of https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.hi https://huggingface.co/datasets/xtreme https://huggingface.co/datasets/xquad https://huggingface.co/datasets/databricks/databricks-dolly-15k/viewer/default/train?p=17&f[category][value]=%27closed_qa%27 (closed-qa only) textquestion-answering10K<n<100K0 likes31 downloads3y agoHugging Face13Speech-data /Hindi-Speech-Dataset 🎧 Hindi Speech Dataset The Hindi Speech Dataset is a high-quality and structured speech audio dataset developed to support modern AI systems that rely on diverse audio data and scalable voice data. It contains 132 hours of recordings distributed across 565 files, available in MP3 and WAV formats, with a total size of 101 MB. This carefully curated audio dataset provides balanced speaker representation with 49% female and 51% male contributors, covering an age range from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hindi-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes30 downloads6mo agoHugging Face14miraiminds /function-calling-dataset-Hindi-englishtext1K<n<10K0 likes28 downloads2y agoHugging Face15ud-nlp /hindi-speech-recognition-dataset Hindi Telephone Dialogues Dataset - 760 Hours Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes28 downloads8mo agoHugging Face16ganeshjcs /hindi-headline-article-generation Summary hindi-headline-article-generation is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-headline-article-generation.texttext-generation100K<n<1M1 likes26 downloads3y agoHugging Face17ReySajju742 /Hindi-Poetry-Dataset Hindi Transliteration of Urdu Poetry Dataset Welcome to the Hindi Transliteration of Urdu Poetry Dataset! This dataset features Hindi transliterations of traditional Urdu poetry. Each entry in the dataset includes two columns: Title: The transliterated title of the poem in Hindi. Poem: The transliterated text of the Urdu poem rendered in Hindi script. This dataset is perfect for researchers and developers working on cross-script language processing, transliteration models, and… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/Hindi-Poetry-Dataset.texttext-classification1K<n<10K0 likes26 downloads2y agoHugging Face18UniDataPro /hindi-speech-recognition-dataset Hindi Speech Dataset for recognition task Dataset comprises 760 hours of telephone dialogues in Hindi, collected from 1,000+ native speakers across various topics and domains. This dataset boasts an impressive 95% sentence accuracy rate, making it a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K1 likes26 downloads1mo agoHugging Face19akashuee /hindi_visual_genometabular10K<n<100K0 likes26 downloads1y agoHugging Face20CodeWithSomesh /english-hindi-vocab-flashcardsgatedtexttext-classification1K<n<10K2 likes25 downloads1y agoHugging Face21VishalMysore /Hindi_Mithai Dataset Card for Indian Sweets textn<1K0 likes23 downloads3y agoHugging Face22mlexplorer008 /hin_dialect_classificationtext1K<n<10K0 likes21 downloads2y agoHugging Face23d0r1h /HindiNewSummaries How to use this dataset # You can load data using following code and then split into train and validation set from datasets import load_dataset data = load_dataset("d0r1h/HindiNewSummaries") Licensing Note: This license applies to the dataset curation and summaries. Original article copyrights remain with their publishers. This dataset contains news articles and summaries collected from publicly available news websites. The original copyright of the article text… See the full description on the dataset page: https://huggingface.co/datasets/d0r1h/HindiNewSummaries.textsummarization100K<n<1M0 likes21 downloads4mo agoHugging Face24iam-tsr /hindi-sentimentslabel: {'Neutral': 0, 'Positive': 1, 'Negative': 2} texttext-classification100K<n<1M0 likes21 downloads7mo agoHugging Face25shivam9980 /inshorts-hinditext10K<n<100K0 likes20 downloads3y agoHugging Face26bajpaideeksha /english-hindi-colloquial-datasetA curated dataset of colloquial English phrases and their corresponding Hindi translations. This dataset focuses on informal language, including slang, idioms, and everyday expressions, making it ideal for training models that handle casual conversations. Dataset Details: Size:e.g., 500+ phrase pairs] Source: Collected from publicly available conversational datasets, social media, and crowdsourced contributions. Language Pair: English → Hindi Annotations: Each phrase pair is manually verified… See the full description on the dataset page: https://huggingface.co/datasets/bajpaideeksha/english-hindi-colloquial-dataset.texttranslationn<1K2 likes20 downloads2y agoHugging Face27guneetsk99 /hindi_instruction_set_187Ktext100K<n<1M2 likes19 downloads3y agoHugging Face28wong132 /bengali-hindi-number-blindspot Blind Spots of Frontier Models: Bengali & Hindi Number Word-to-Digit Conversion Summary This dataset documents a critical blind spot in small open-source language models: failure to correctly convert Bengali and Hindi number words into their digit equivalents. Bengali and Hindi share the South Asian number system (hazar/হাজার, lakh/লাখ, crore/কোটি), and all three tested models consistently fail at this fundamental conversion step. IMPORTANT: Arithmetic calculation errors… See the full description on the dataset page: https://huggingface.co/datasets/wong132/bengali-hindi-number-blindspot.texttext-generationn<1K0 likes19 downloads7mo agoHugging Face29TokenBender /sentence_retrieval_hindi_SFTtext10K<n<100K2 likes18 downloads3y agoHugging Face30TokenBender /e5_FT_sentence_retrieval_task_Hinditext10K<n<100K0 likes17 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.