datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SQuAD_HindiThis dataset is created by translating a part of the Stanford QA dataset.
It contains 5k QA pairs from the original SQuad dataset translated to Hindi using the googletrans api.
hatecheck-hindi
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-hindi.HinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/HinGE.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.hinemoIndian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.hindi_visual_genomemindbridge-phq9-hindi-seeds
MindBridge Hindi PHQ-9/GAD-7 — Gold Seeds (144 rows)
Hand-authored Hindi seeds for PHQ-9 + GAD-7 screening across three personas
(postnatal_mother, older_woman, man) in 1:1:1 distribution. Authored via
SuperWhisper Scribe with cloud LLM post-process; all rows
human-reviewed with review_status=accepted.
This seed set drives Phase B teacher expansion (in-context exemplars for
Gemma 4 26B-A4B MoE on Vertex MaaS) plus 24 Item-9 (suicidality) extras
authored separately. See companion… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-seeds.google_go_emotions_hindi_translatedHinGEAbstract
Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately, text generation and evaluation are relatively understudied due to the scarcity of high-quality resources in code-mixed languages where the words and phrases from multiple languages are mixed in a single utterance of text and speech. To address this… See the full description on the dataset page: https://huggingface.co/datasets/ShreyaDev/HinGE.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/chinna887/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/Vedantsc-1110/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.longer-context-hindi-2e5_FT_sentence_retrieval_task_Hindi_miniHindi_englishscienceqa-multilingual-hindi#ScienceQA Hindi Translation Dataset
##Dataset Description
This dataset is a Hindi-translated version of the original ScienceQA dataset. It includes multiple-choice science questions, with fields for:
Images (optional visual context),
Hints (optional support text),
English questions and their Hindi translations,
Multiple answer choices,
Correct answers.
This translation is intended to support multilingual education research, question-answering in Hindi, and fairness studies in multilingual… See the full description on the dataset page: https://huggingface.co/datasets/model2me/scienceqa-multilingual-hindi.CIC-MalMem-2022CIC-MalMem-2022 was created by researchers at the Canadian Institute for Cybersecurity (CIC) at the University of New Brunswick.
The details are described on the website https://www.unb.ca/cic/datasets/malmem-2022.html, and in their paper mentioned on that site:
Tristan Carrier, Princy Victor, Ali Tekeoglu, Arash Habibi Lashkari,” Detecting Obfuscated Malware using Memory Feature Engineering”,
The 8th International Conference on Information Systems Security and Privacy (ICISSP), 2022.
The… See the full description on the dataset page: https://huggingface.co/datasets/hina-Code/CIC-MalMem-2022.Indian-EN-HIN-Dataste-TTSchungjudam
