CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkurman /hindawi-journals-2007-2023 Hindawi Academic Papers Dataset (CC BY 4.0 Compatible) Dataset Description This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content. Dataset Summary Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.texttext-generation100K<n<1M5 likes601 downloads1y agoHugging Face02kaifahmad /indian-history-hindi-QA-3.4k Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This dataset contains 3.47k top-notch question-answer pairs about Indian History in Hindi. Curated by: Mohd Kaif Language(s) (NLP): Hindi License: apache-2.0 textquestion-answering1K<n<10K0 likes180 downloads3y agoHugging Face03HINT-lab /DeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration. We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.tabularquestion-answering100K<n<1M0 likes168 downloads2y agoHugging Face04inverse-scaling /hindsight-neglect-10shot inverse-scaling/hindsight-neglect-10shot (‘The Floating Droid’) General description This task tests whether language models are able to assess whether a bet was worth taking based on its expected value. The author provides few shot examples in which the model predicts whether a bet is worthwhile by correctly answering yes or no when the expected value of the bet is positive (where the model should respond that ‘yes’, taking the bet is the right decision) or negative (‘no’… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/hindsight-neglect-10shot.textmultiple-choicen<1K5 likes138 downloads4y agoHugging Face05HINT-lab /Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration. 🔥 Updates [2025-3-3]: We released our paper. [2025-2-25]: We released our codes, models and datasets. 🏴󠁶󠁵󠁭󠁡󠁰󠁿 Overview We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.tabularquestion-answering100K<n<1M0 likes127 downloads2y agoHugging Face06shreyansh12183 /shreyansh-hinglish-english-stem-500k 🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations. 📖 Overview In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.textquestion-answering100K<n<1M0 likes106 downloads5d agoHugging Face07araag2 /HINT HINT: Hierarchical interaction network for clinical-trial-outcome predictions Dataset Description Links Homepage: Github.io Repository: Github Paper: arXiv Contact (Original Authors): Tianfan Fu (futianfan@gmail.com) Contact (Curator):Artur Guimarães (artur.guimas@gmail.com) Dataset Summary Clinical trials are crucial for drug development but are time consuming, expensive, and often burdensome on patients. More importantly, clinical… See the full description on the dataset page: https://huggingface.co/datasets/araag2/HINT.textquestion-answering10K<n<100K0 likes87 downloads11mo agoHugging Face08nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes86 downloads3y agoHugging Face09rmahesh /UP_CET_Hindi_examstextquestion-answering1K<n<10K0 likes77 downloads2y agoHugging Face10InfoBayAI /Hindi-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of Hindi STEM Question Answering (QA) data, containing 1,854,832 question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, reasoning, problem-solving, and educational learning in Hindi. The dataset consists of multiple-choice question answering (MCQA) samples across core STEM domains including Physics, Mathematics, Chemistry, Biology, and General… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K0 likes54 downloads8d agoHugging Face11OdiaGenAI /instruction_set_hindi_1035The dataset has been created using OliveFarm web application. Following domains have been covered in this dataset:- Art Sports (Cricket, Football, Olympics) Politics History Cooking Environment Music Contributors: - Shahid Parul. textquestion-answering1K<n<10K1 likes52 downloads3y agoHugging Face12me-nabi /hindikrishi-farmer-advisory-dataset 🌾 HindiKrishi — Farmer Advisory Dataset 21,069 instruction-response pairs for training agricultural crop advisory models in Hindi and English, grounded in ICAR guidelines. Dataset Details Detail Value Total Examples 21,069 Languages Hindi (primary), English Format JSONL (instruction, input, output) Domain Indian agriculture — crop diseases, pesticides, fertilizers, schemes License Apache 2.0 Format Each example follows the… See the full description on the dataset page: https://huggingface.co/datasets/me-nabi/hindikrishi-farmer-advisory-dataset.texttext-generation10K<n<100K0 likes45 downloads2mo agoHugging Face13UnfaithRL /mmlu_hinted_questions MMLU Hinted Questions Dataset Description This dataset contains multiple-choice questions derived from MMLU and augmented with misleading hints. The misleading hints are intentionally designed to point to an incorrect answer. The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards. Specifically, it was used to investigate whether language models follow… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/mmlu_hinted_questions.tabularquestion-answering10K<n<100K0 likes40 downloads3mo agoHugging Face14JamshidJDMY /HintQA HintQA: Exploring Hint Generation Approaches in Open-Domain Question Answering HintQA revolutionizes the field of automatic question answering by introducing a novel context preparation method that utilizes Automatic Hint Generation. Unlike traditional QA systems that rely on either retrieval-based methods (sourcing documents from databases like Wikipedia) or generation-based approaches (using large language models to generate context), HintQA prompts large language models to… See the full description on the dataset page: https://huggingface.co/datasets/JamshidJDMY/HintQA.question-answering0 likes36 downloads2y agoHugging Face15alonmiron /mmlu_hinted_huggingfaceThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.question-answering0 likes33 downloads2y agoHugging Face16InfoBayAI /Hindi-Non-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of Hindi Non-STEM Question Answering (QA) data, containing over 1.4 million question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, knowledge retrieval, and educational learning in Hindi. It is part of a broader collection of 6.5+ million question-answer pairs spanning multiple languages and domains. The dataset consists of multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Hindi-Non-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K1 likes31 downloads10d agoHugging Face17maharnab /hindi_instructThis dataset was created for the "Unlock Global Communication with Gemma" competition on Kaggle. It combines multiple datasets to capture a diverse range of topics and use cases: OdiaGenAI/instruction_set_hindi_1035: Includes instructions and responses related to art, culture, history, cooking, environment, music, and sports. SherryT997/HelpSteer-hindi: Focuses on general question-answering conversations. kaifahmad/indian-history-hindi-QA-3.4k: Contains questions and answers specifically… See the full description on the dataset page: https://huggingface.co/datasets/maharnab/hindi_instruct.textquestion-answering1K<n<10K1 likes30 downloads2y agoHugging Face18Khyatimirani /pcos_question_answer_hindi PCOS Hindi Lifestyle & Clinical Q&A Dataset Dataset Details Dataset Description This dataset contains patient-facing conversational question–answer pairs in Hindi (Devanagari script) focused on Polycystic Ovary Syndrome (PCOS/PCOD). The dataset is designed to support training and evaluation of healthcare conversational AI systems that provide lifestyle and general clinical guidance for women diagnosed with PCOS. All conversations are structured in a chat format… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos_question_answer_hindi.textquestion-answeringn<1K0 likes28 downloads7mo agoHugging Face19SherryT997 /HelpSteer-hinditexttext-classification1K<n<10K1 likes27 downloads3y agoHugging Face20bingbangboom /fka-awesome-chatgpt-prompts-hindi🧠 Awesome ChatGPT Prompts in Hindi [CSV dataset] This is a hindi translated dataset repository of Awesome ChatGPT Prompts fka/awesome-chatgpt-prompts View All Original Prompts on GitHub textquestion-answeringn<1K1 likes26 downloads2y agoHugging Face21cmeraki /hindi_eval_general_mcqtextquestion-answering1K<n<10K3 likes25 downloads3y agoHugging Face22CodeWithSomesh /english-hindi-vocab-flashcardsgatedtexttext-classification1K<n<10K2 likes25 downloads1y agoHugging Face23OdiaGenAI /roleplay_hindiThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Hindi using OdiaGenAI English=>Indic translation app. textquestion-answering1K<n<10K1 likes24 downloads3y agoHugging Face24OdiaGenAI /health_hindi_200Contributors: - Sonal Khosla textquestion-answeringn<1K1 likes24 downloads3y agoHugging Face25Process-Venue /Hindi-Marathi-Synonyms Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह) Overview This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications. The dataset provides word-synonym pairs that can be used for tasks like: Semantic analysis Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.texttext-classification1K<n<10K0 likes24 downloads2y agoHugging Face26HydraIndicLM /Hindi_Train_ClosedDomainQAThe dataset is the Hindi-only and processed version of https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.hi https://huggingface.co/datasets/xtreme https://huggingface.co/datasets/xquad https://huggingface.co/datasets/databricks/databricks-dolly-15k/viewer/default/train?p=17&f[category][value]=%27closed_qa%27 (closed-qa only) textquestion-answering10K<n<100K0 likes21 downloads3y agoHugging Face27harshraj /dolly15k_hinglish_dataset_cleanedtextquestion-answering10K<n<100K2 likes21 downloads2y agoHugging Face28damerajee /ShareGPT4V-hin Dataset details It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Multi-Modal Models (LMMs) during both the pre-training and supervised fine-tuning stages. This advancement aims to bring LMMs towards GPT4-Vision capabilities. sharegpt4v_instruct_gpt4-vision_cap100k.json is generated by GPT4-Vision (ShareGPT4V). This dataset is Hindi-translated version of the ShareGPT4V This dataset is intended only for Fine-tuning The images can be… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/ShareGPT4V-hin.textvisual-question-answering100K<n<1M0 likes18 downloads2y agoHugging Face29damerajee /Hindi-LLaVA-CC3M-Pretrain-595K LLaVA Visual Instruct CC3M 595K Pretrain Dataset Card Dataset details Dataset type: LLaVA Visual Instruct CC3M Pretrain 595K is a subset of CC-3M dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. We aim to build large multimodal towards GPT-4 vision/language capability. Dataset date:… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/Hindi-LLaVA-CC3M-Pretrain-595K.textvisual-question-answering100K<n<1M0 likes16 downloads2y agoHugging Face30QuantumMik /alpaca_hindi_small Alpaca Hindi Small This is a synthesized dataset created by translation of alpaca dataset from English to Hindi language. textquestion-answering1K<n<10K1 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.