datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindawi-journals-2007-2023
Hindawi Academic Papers Dataset (CC BY 4.0 Compatible)
Dataset Description
This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content.
Dataset Summary
Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.indian-history-hindi-QA-3.4k
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This dataset contains 3.47k top-notch question-answer pairs about Indian History in Hindi.
Curated by: Mohd Kaif
Language(s) (NLP): Hindi
License: apache-2.0
DeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration.
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.hindsight-neglect-10shot
inverse-scaling/hindsight-neglect-10shot (‘The Floating Droid’)
General description
This task tests whether language models are able to assess whether a bet was worth taking based on its expected value. The author provides few shot examples in which the model predicts whether a bet is worthwhile by correctly answering yes or no when the expected value of the bet is positive (where the model should respond that ‘yes’, taking the bet is the right decision) or negative (‘no’… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/hindsight-neglect-10shot.Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration.
🔥 Updates
[2025-3-3]: We released our paper.
[2025-2-25]: We released our codes, models and datasets.
🏴 Overview
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.HINT
HINT: Hierarchical interaction network for clinical-trial-outcome predictions
Dataset Description
Links
Homepage:
Github.io
Repository:
Github
Paper:
arXiv
Contact (Original Authors):
Tianfan Fu (futianfan@gmail.com)
Contact (Curator):Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
Clinical trials are crucial for drug development but are time consuming, expensive, and often burdensome on patients. More importantly, clinical… See the full description on the dataset page: https://huggingface.co/datasets/araag2/HINT.chaii-hindi-and-tamil-question-answeringUP_CET_Hindi_examsHindi-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi STEM Question Answering (QA) data, containing 1,854,832 question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, reasoning, problem-solving, and educational learning in Hindi.
The dataset consists of multiple-choice question answering (MCQA) samples across core STEM domains including Physics, Mathematics, Chemistry, Biology, and General… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-STEM-QA-MCQ-Dataset.instruction_set_hindi_1035The dataset has been created using OliveFarm web application.
Following domains have been covered in this dataset:-
Art
Sports (Cricket, Football, Olympics)
Politics
History
Cooking
Environment
Music
Contributors: -
Shahid
Parul.
hindikrishi-farmer-advisory-dataset
🌾 HindiKrishi — Farmer Advisory Dataset
21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines.
Dataset Details
Detail
Value
Total Examples
21,069
Languages
Hindi (primary), English
Format
JSONL (instruction, input, output)
Domain
Indian agriculture — crop diseases, pesticides, fertilizers, schemes
License
Apache 2.0
Format
Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.mmlu_hinted_questions
MMLU Hinted Questions
Dataset Description
This dataset contains multiple-choice questions derived from MMLU and augmented with misleading hints. The misleading hints are intentionally designed to point to an incorrect answer.
The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards.
Specifically, it was used to investigate whether language models follow… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/mmlu_hinted_questions.HintQA
HintQA: Exploring Hint Generation Approaches in Open-Domain Question Answering
HintQA revolutionizes the field of automatic question answering by introducing a novel context preparation method that utilizes Automatic Hint Generation. Unlike traditional QA systems that rely on either retrieval-based methods (sourcing documents from databases like Wikipedia) or generation-based approaches (using large language models to generate context), HintQA prompts large language models to… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/HintQA.mmlu_hinted_huggingfaceThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.Hindi-Non-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains.
The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.hindi_instructThis dataset was created for the "Unlock Global Communication with Gemma" competition on Kaggle. It combines multiple datasets to capture a diverse range of topics and use cases:
OdiaGenAI/instruction_set_hindi_1035: Includes instructions and responses related to art, culture, history, cooking, environment, music, and sports.
SherryT997/HelpSteer-hindi: Focuses on general question-answering conversations.
kaifahmad/indian-history-hindi-QA-3.4k: Contains questions and answers specifically… See the full description on the dataset page: https://huggingface.co/datasets/maharnab/hindi_instruct.pcos_question_answer_hindi
PCOS Hindi Lifestyle & Clinical Q&A Dataset
Dataset Details
Dataset Description
This dataset contains patient-facing conversational question–answer pairs in Hindi (Devanagari script) focused on Polycystic Ovary Syndrome (PCOS/PCOD).
The dataset is designed to support training and evaluation of healthcare conversational AI systems that provide lifestyle and general clinical guidance for women diagnosed with PCOS.
All conversations are structured in a chat format… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos_question_answer_hindi.HelpSteer-hindifka-awesome-chatgpt-prompts-hindi🧠 Awesome ChatGPT Prompts in Hindi [CSV dataset]
This is a hindi translated dataset repository of Awesome ChatGPT Prompts fka/awesome-chatgpt-prompts
View All Original Prompts on GitHub
hindi_eval_general_mcqenglish-hindi-vocab-flashcardsroleplay_hindiThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Hindi using OdiaGenAI English=>Indic translation app.
health_hindi_200Contributors: -
Sonal Khosla
Hindi-Marathi-Synonyms
Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह)
Overview
This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications.
The dataset provides word-synonym pairs that can be used for tasks like:
Semantic analysis
Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.Hindi_Train_ClosedDomainQAThe dataset is the Hindi-only and processed version of
https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.hi
https://huggingface.co/datasets/xtreme
https://huggingface.co/datasets/xquad
https://huggingface.co/datasets/databricks/databricks-dolly-15k/viewer/default/train?p=17&f[category][value]=%27closed_qa%27 (closed-qa only)
dolly15k_hinglish_dataset_cleanedShareGPT4V-hin
Dataset details
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Multi-Modal Models (LMMs) during both the pre-training and supervised fine-tuning stages. This advancement aims to bring LMMs towards GPT4-Vision capabilities.
sharegpt4v_instruct_gpt4-vision_cap100k.json is generated by GPT4-Vision (ShareGPT4V).
This dataset is Hindi-translated version of the ShareGPT4V
This dataset is intended only for Fine-tuning
The images can be… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/ShareGPT4V-hin.Hindi-LLaVA-CC3M-Pretrain-595K
LLaVA Visual Instruct CC3M 595K Pretrain Dataset Card
Dataset details
Dataset type:
LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution.
Captions are also associated with BLIP synthetic caption for reference.
It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
We aim to build large multimodal towards GPT-4 vision/language capability.
Dataset date:… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/Hindi-LLaVA-CC3M-Pretrain-595K.alpaca_hindi_small
Alpaca Hindi Small
This is a synthesized dataset created by translation of alpaca dataset from English to Hindi language.
