datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.Hinglish_Dataset_instruction_and_rawprism-hinglish-hate-speech
PRISM - Code-Mixed Hinglish Hate-Speech Dataset
Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project
Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text
(RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle.
Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track
Summary
Attribute
Value
Total samples (raw)
29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.scope-benchmark
SCOPE Benchmark
Evaluation benchmark for the HRI '26 paper SCOPE: A Real-Time Natural Language Camera Agent at the Edge (arXiv:2606.02951). Test-only — no train split. 541 questions × 4 Blender scenes × 8 task categories.
The code that runs this benchmark lives at github.com/HindsboNikolaj/SCOPE.
When you chain a language model and a vision model together, how do you know which one failed?
Contents
scope-benchmark/
scope_541.csv… See the full description on the dataset page: https://huggingface.co/datasets/HindsboNikolaj/scope-benchmark.romanized_hindi
Romanized Hindi Dataset
Dataset Description
The Romanized Hindi Dataset is a collection of Hindi text paired with its Romanized (Latin script) representation.
It has been created by combining multiple sources, including open datasets, synthetic generation, and rule-based transliteration methods.
The dataset is designed for training and evaluating Hindi↔Roman transliteration models.
Language(s): Hindi, Romanized Hindi
Size: ~1.82M rows
License: MIT (check with source… See the full description on the dataset page: https://huggingface.co/datasets/sk-community/romanized_hindi.SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset.
It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api.
hatecheck-hindi
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.HinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.hindi-english-code-mixed-tweets-sentimentchatbot_arena_conversations_hinglishThe dataset is created by translating "lmsys/chatbot_arena_conversations" dataset.
link to original datset - https://huggingface.co/datasets/lmsys/chatbot_arena_conversations
Original dataset contain two conversation from model_a and model_b and also given winner model between these two model conversation.
I have selected winner conversation and converted that user query and assistant answer into hinglish language using Gemini pro
hinglish-youtube-sentiments-dataset
Hinglish YouTube Comments Sentiment Dataset
A manually annotated dataset of 3,190 Hinglish YouTube comments for 3-class sentiment classification. Hinglish is the code-mixed Hindi-English language used by hundreds of millions of Indians online — written in Roman script, mixing Hindi and English words fluidly within the same sentence.
This dataset was created because no sufficiently large, cleanly annotated Hinglish sentiment dataset existed for YouTube comment data specifically.… See the full description on the dataset page: https://huggingface.co/datasets/shae2977/hinglish-youtube-sentiments-dataset.HinduTamil-News-Articles-Dataset
HinduTamil News Articles Dataset
Overview
This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text.
Motivation
This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes.
Data Sources and collection method
The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Bobble-Hinglish-Sports-Dataset_BHSD
Dataset Card for Bobble Hinglish Sports Dataset (BHSD)
Dataset Description
The Bobble Hinglish Sports Dataset is a meticulously curated collection of 7,029 code-mixed sentences spanning various sports categories. It includes human annotations across seven distinct sports categories, along with out-of-scope values. This testing dataset is specifically designed to enhance NLP models' ability to understand Hinglish sports content, making it valuable for intent detection… See the full description on the dataset page: https://huggingface.co/datasets/BobbleAI/Bobble-Hinglish-Sports-Dataset_BHSD.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.hinemoHinglish-Hindi-Transliteration-Dataset
Hinglish-Hindi Transliteration Dataset
We are pleased to release this unique dataset focused on transliteration between Hinglish (Hindi written in Roman script) and Devanagari Hindi. This dataset aims to address the limitations of current models in accurately transliterating words and phrases as they are commonly used, preserving their original form and meaning. Unlike translation datasets, this resource focuses on phonetic equivalence rather than semantic transformation. For… See the full description on the dataset page: https://huggingface.co/datasets/codebyam/Hinglish-Hindi-Transliteration-Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Movie_Review_Sentiment_HindiHindi_sentimentHindi-Non-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains.
The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.hindi-speech-recognition-dataset
Hindi Telephone Dialogues Dataset - 760 Hours
Dataset comprises 760 hours of high-quality audio recordings from 1,000+ native Hindi speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating Hindi speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in Hindi for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/hindi-speech-recognition-dataset.HindiHindi-Speech-Dataset
🎧 Hindi Speech Dataset
The Hindi Speech Dataset is a high-quality and structured speech audio dataset developed to support modern AI systems that rely on diverse audio data and scalable voice data. It contains 132 hours of recordings distributed across 565 files, available in MP3 and WAV formats, with a total size of 101 MB. This carefully curated audio dataset provides balanced speaker representation with 49% female and 51% male contributors, covering an age range from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Hindi-Speech-Dataset.function-calling-dataset-Hindi-englishhindi_visual_genomeHinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/VaniAgent/Hinglish-Everyday-Conversations-1M.
