CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes353 downloads2y agoHugging Face02Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes129 downloads9mo agoHugging Face03pankajbiswas6 /prism-hinglish-hate-speech PRISM - Code-Mixed Hinglish Hate-Speech Dataset Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text (RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle. Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track Summary Attribute Value Total samples (raw) 29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.texttext-classification10K<n<100K0 likes128 downloads3mo agoHugging Face04HindsboNikolaj /scope-benchmark SCOPE Benchmark Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories. The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE. When you chain a language model and a vision model together, how do you know which one failed? Contents scope-benchmark/ scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.documentvisual-question-answeringn<1K0 likes126 downloads4mo agoHugging Face05sk-community /romanized_hindi Romanized Hindi Dataset Dataset Description The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation. It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods. The dataset is designed for training and evaluating Hindi↔Roman transliteration models. Language(s): Hindi, Romanized Hindi Size: ~1.82M rows License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.text1M<n<10M0 likes115 downloads1y agoHugging Face06aneesh-b /SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset. It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api. tabular1K<n<10K0 likes101 downloads4y agoHugging Face07Paul /hatecheck-hindi Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.tabulartext-classification1K<n<10K1 likes98 downloads4y agoHugging Face08LingoIITGN /HinGEAbstract Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.tabulartranslation1K<n<10K1 likes92 downloads2y agoHugging Face09ysangam /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K1 likes86 downloads4mo agoHugging Face10Abhishek4896 /hindi-english-code-mixed-tweets-sentimenttexttext-classificationn<1K0 likes82 downloads1y agoHugging Face11one-thing /chatbot_arena_conversations_hinglishThe dataset is created by translating "lmsys/chatbot_arena_conversations" dataset. link to original datset - https://huggingface.co/datasets/lmsys/chatbot_arena_conversations Original dataset contain two conversation from model_a and model_b and also given winner model between these two model conversation. I have selected winner conversation and converted that user query and assistant answer into hinglish language using Gemini pro text10K<n<100K5 likes67 downloads3y agoHugging Face12shae2977 /hinglish-youtube-sentiments-dataset Hinglish YouTube Comments Sentiment Dataset A manually annotated dataset of 3,190 Hinglish YouTube comments for 3-class sentiment classification. Hinglish is the code-mixed Hindi-English language used by hundreds of millions of Indians online — written in Roman script, mixing Hindi and English words fluidly within the same sentence. This dataset was created because no sufficiently large, cleanly annotated Hinglish sentiment dataset existed for YouTube comment data specifically.… See the full description on the dataset page: https://huggingface.co/datasets/shae2977/hinglish-youtube-sentiments-dataset.text1K<n<10K0 likes64 downloads4mo agoHugging Face13Shwetasss /HinduTamil-News-Articles-Dataset HinduTamil News Articles Dataset Overview This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text. Motivation This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes. Data Sources and collection method The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.texttext-classification10K<n<100K1 likes61 downloads3y agoHugging Face14DatasetNewUser /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K0 likes61 downloads15d agoHugging Face15BobbleAI /Bobble-Hinglish-Sports-Dataset_BHSD Dataset Card for Bobble Hinglish Sports Dataset (BHSD) Dataset Description The Bobble Hinglish Sports Dataset is a meticulously curated collection of 7,029 code-mixed sentences spanning various sports categories. It includes human annotations across seven distinct sports categories, along with out-of-scope values. This testing dataset is specifically designed to enhance NLP models' ability to understand Hinglish sports content, making it valuable for intent detection… See the full description on the dataset page: https://huggingface.co/datasets/BobbleAI/Bobble-Hinglish-Sports-Dataset_BHSD.texttoken-classification1K<n<10K0 likes57 downloads1y agoHugging Face16ganeshjcs /hindi-article-summarization Summary hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.texttext-generation10K<n<100K0 likes56 downloads3y agoHugging Face17theanonymous524 /hinemotabular10K<n<100K1 likes51 downloads22d agoHugging Face18codebyam /Hinglish-Hindi-Transliteration-Datasetgated Hinglish-Hindi Transliteration Dataset We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.texttext-generation1K<n<10K2 likes43 downloads1y agoHugging Face19harshlimkar /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K0 likes43 downloads29d agoHugging Face20szjkdsldf /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K0 likes42 downloads23d agoHugging Face21arka15 /Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset Indian Scam Communication Dataset Overview The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India. The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety. Motivation India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.tabulartext-classification10K<n<100K0 likes41 downloads26d agoHugging Face22Process-Venue /Movie_Review_Sentiment_Hinditexttext-classification1K<n<10K0 likes37 downloads10mo agoHugging Face23sepidmnorozy /Hindi_sentimenttextn<1K1 likes32 downloads4y agoHugging Face24InfoBayAI /Hindi-Non-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains. The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K1 likes31 downloads9d agoHugging Face25ud-nlp /hindi-speech-recognition-dataset Hindi Telephone Dialogues Dataset - 760 Hours Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.textautomatic-speech-recognitionn<1K0 likes30 downloads8mo agoHugging Face26harshitkaran /Hinditext100K<n<1M3 likes29 downloads4y agoHugging Face27Speech-data /Hindi-Speech-Dataset 🎧 Hindi Speech Dataset The Hindi Speech Dataset is a high-quality and structured speech audio dataset developed to support modern AI systems that rely on diverse audio data and scalable voice data. It contains 132 hours of recordings distributed across 565 files, available in MP3 and WAV formats, with a total size of 101 MB. This carefully curated audio dataset provides balanced speaker representation with 49% female and 51% male contributors, covering an age range from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hindi-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes28 downloads6mo agoHugging Face28miraiminds /function-calling-dataset-Hindi-englishtext1K<n<10K0 likes27 downloads2y agoHugging Face29akashuee /hindi_visual_genometabular10K<n<100K0 likes27 downloads1y agoHugging Face30VaniAgent /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/VaniAgent/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M1 likes27 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.