CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdaMLLab /HinMix HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication. We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.texttext-generation100M<n<1B1 likes1.7k downloads8mo agoHugging Face02mkurman /hindawi-journals-2007-2023 Hindawi Academic Papers Dataset (CC BY 4.0 Compatible) Dataset Description This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content. Dataset Summary Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.texttext-generation100K<n<1M5 likes601 downloads1y agoHugging Face03Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes353 downloads2y agoHugging Face04zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes337 downloads3y agoHugging Face05damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes222 downloads2y agoHugging Face06findnitai /english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC source : 1 - Human Annotated , source : 0 - Synthetically Generated texttranslation100K<n<1M23 likes180 downloads3y agoHugging Face07alielfilali01 /Hindawi-Books-dataset Dataset Card for "Hindawi Books Dataset" Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic. Dataset Description Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.texttext-generation10K<n<100K14 likes158 downloads3y agoHugging Face08zicsx /OSCAR-2301-Hindi-Cleaned Dataset Card for "OSCAR-2301-Hindi-Cleaned-2.0" More Information needed texttext-generation100K<n<1M0 likes148 downloads3y agoHugging Face09Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes129 downloads9mo agoHugging Face10Lots-of-LoRAs /task1009_pib_translation_bengali_hindi Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1009_pib_translation_bengali_hindi Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1009_pib_translation_bengali_hindi.texttext-generation1K<n<10K0 likes120 downloads2y agoHugging Face11HINT-lab /Qwen2.5-7B-Instruct-Self-Calibration Efficient Test-Time Scaling via Self-Calibration This repository contains datasets used in the paper Efficient Test-Time Scaling via Self-Calibration. The datasets are used to evaluate the effectiveness of test-time scaling methods for LLMs. Each config_name in the metadata refers to a different reasoning dataset. More detailed descriptions of each dataset are needed. Consider adding a section for each config_name with a description, statistics, and any other relevant… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Qwen2.5-7B-Instruct-Self-Calibration.tabulartext-generation100K<n<1M0 likes119 downloads2y agoHugging Face12Lots-of-LoRAs /task427_hindienglish_corpora_hi-en_language_identification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task427_hindienglish_corpora_hi-en_language_identification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task427_hindienglish_corpora_hi-en_language_identification.texttext-generation1K<n<10K0 likes112 downloads2y agoHugging Face13rvv-karma /English-Hinglish-TOP English Hinglish (TOP Dataset) This dataset is generated from Hinglish-TOP Dataset. Data distribution: Train a. Human Generated - 6513 b. Synthetically generated - 170083 Validation a. Human Generated - 1390 b. Synthetically generated - 0 Test a. Human Generated - 6513 b. Synthetically generated - 0 texttranslation100K<n<1M1 likes109 downloads3y agoHugging Face14shreyansh12183 /shreyansh-hinglish-english-stem-500k 🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations. 📖 Overview In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.textquestion-answering100K<n<1M0 likes106 downloads4d agoHugging Face15Sujalvc /hinglish-instruct-dataset Akshar Hinglish Instruct Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns. 1. Dataset Overview Total Examples: 10,378 Base Set: 9,999 instruction-following pairs Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.texttext-generation10K<n<100K1 likes97 downloads4mo agoHugging Face16ar0cket1 /hintedselfteacher-nemotron-math-v2-AoPS hintedselfteacher-nemotron-math-v2-AoPS This dataset contains a training-ready hinted self-teacher split derived from the AoPS split of nvidia/Nemotron-Math-v2. The source problems were filtered to the AoPS split with the medium/notool solve rate between 2 and 6. Hints were generated with GPT-5.5 medium using an h17_nt hint-generation prompt. This hint type was close to the best hint type found after doing hint mutations, based on qualitative analysis of token-level hinted… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/hintedselfteacher-nemotron-math-v2-AoPS.tabulartext-generation10K<n<100K1 likes79 downloads3mo agoHugging Face17shiv96 /mmlu-hint-faithfulness-traces MMLU and GPQA hint faithfulness traces Model-specific configurations Configuration Rows Model and content qwen3-8b 3,486 Existing Qwen3-8B generations and judgments qwen3-8b_answer_changes 1,137 Existing Qwen3-8B valid answer changes gpt-5.6-luna 3,486 Luna final responses, observable reasoning summaries and judgments gpt-5.6-luna_answer_changes 966 Luna valid answer changes with the same labels All sets contain MMLU and GPQA and use a test… See the full description on the dataset page: https://huggingface.co/datasets/shiv96/mmlu-hint-faithfulness-traces.texttext-generation1K<n<10K0 likes79 downloads8d agoHugging Face18rvv-karma /English-Hinglish English Hinglish English to Hinglish Dataset processed from findnitai/english-to-hinglish. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC texttranslation100K<n<1M0 likes77 downloads3y agoHugging Face19nis12ram /nisram-hindi-text-0.0 Dataset Card for nisram-hindi-text-0.0 Quick Summary Language: Hindi Size: ~618,000 raw text samples (206k from each source, before deduplication) and 602,000 after deduplication Content: Short- to medium-length Hindi text (up to ~5000 characters), drawn from web-crawled corpora Fields: One field text (string) containing the Hindi text Sources: mC4-Hindi-Cleaned-3.0 OSCAR-2301-Hindi-Cleaned-2.0 ai4bharat/sangraha (Hindi, verified) Deduplication: Performed… See the full description on the dataset page: https://huggingface.co/datasets/nis12ram/nisram-hindi-text-0.0.texttext-generation100K<n<1M0 likes76 downloads1y agoHugging Face20Lots-of-LoRAs /task1204_atomic_classification_hinderedby Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1204_atomic_classification_hinderedby Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1204_atomic_classification_hinderedby.texttext-generation1K<n<10K0 likes73 downloads2y agoHugging Face21vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes73 downloads16d agoHugging Face22smangrul /hinglish_self_instruct_v0 Hinglish Instruct Dataset using Self Instruct method The prompt used for generating the samples: You are asked to come up with a set of 50 diverse task instructions in Hinglish or Hindi. These task instructions will be given to a GPT model and we will evaluate the GPT model for completing the instructions. Here are the requirements: 1. Try not to repeat the verb for each instruction to maximize diversity. 2. The language used for the instruction also should be diverse. For example… See the full description on the dataset page: https://huggingface.co/datasets/smangrul/hinglish_self_instruct_v0.texttext-generation1K<n<10K8 likes72 downloads3y agoHugging Face23thinkedgeAI /Hindi-Niband Dataset Name: Hindi- Niband (Massive Hindi language Text Dataset) Dataset Overview This dataset is a comprehensive collection of text data consisting of more than 10 billion tokens. It encompasses a wide range of sources, including Wikipedia articles, news articles, email transcripts, and generated prompt text. Specific Hindi language data columns have been extracted from the CulturaX dataset, which is a large, cleaned, and multilingual dataset for large language models.… See the full description on the dataset page: https://huggingface.co/datasets/thinkedgeAI/Hindi-Niband.texttext-generation1M<n<10M0 likes58 downloads3y agoHugging Face24Pastaaaaa2003 /Hindi-speech-instructgated Hindi LLaMA-Omni Instruct Dataset A Hindi speech instruction-following dataset designed for training speech-language models such as LLaMA-Omni. Each example pairs a spoken Hindi user question (audio) with a text assistant response. Dataset Summary Property Value Language Hindi (hi) Total examples ~110,718 Train split ~105,000 examples (batches 001–210) Validation split ~5,500 examples (batches 211–222) Audio format FLAC, 16,000 Hz mono… See the full description on the dataset page: https://huggingface.co/datasets/Pastaaaaa2003/Hindi-speech-instruct.audioautomatic-speech-recognition100K<n<1M0 likes57 downloads3mo agoHugging Face25saidutta69 /hinglish-bench Hinglish-Bench 📄 Paper: Hinglish-Bench — Reference-Free Benchmark for LLM Hinglish Text Generation (gist preprint) A reference-free benchmark for measuring how well LLMs generate natural Roman-script Hinglish — the Hindi-English code-mixing that hundreds of millions of Indians actually speak, type, and read online. Reference-free by design. Hinglish has no canonical spelling and no single "correct" rendering, so there are no gold references and no BLEU. Quality is… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/hinglish-bench.texttext-generationn<1K0 likes57 downloads2mo agoHugging Face26ganeshjcs /hindi-article-summarization Summary hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Hindi Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.texttext-generation10K<n<100K0 likes56 downloads3y agoHugging Face27AmareshHebbar /hindi-medical-sft Hindi Medical Reasoning (Medical-o1-SFT) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Medical questions → detailed chain-of-thought reasoning and clinical answers Why download this Fine-tune models for Hindi-language medical Q&A, build ABDM-compatible clinical assistants, or create multilingual medical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/hindi-medical-sft.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face28Transluce /input_ablation_qwen3_8b_mmlu_hint Training Language Models to Explain Their Own Computations (Input Ablations) This dataset is part of the work presented in the paper "Training Language Models to Explain Their Own Computations". It specifically contains data for the Input Ablations task for the Qwen3-8B target model. In this task, explainer models are trained to predict how removing "hint" tokens from an MMLU prompt with a hint changes the output of Qwen3-8B. This helps in understanding the causal relationships… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/input_ablation_qwen3_8b_mmlu_hint.texttext-generation10K<n<100K1 likes51 downloads9mo agoHugging Face29ar0cket1 /qwen3-4b-base-self-distillation-hints Qwen3-4B-Base Self-Distillation Hints 4,481 unique accepted math problems, each with a verifiable target answer, a reference worked solution, and an h17_nt private problem-solving hint. The intended student is Qwen/Qwen3-4B-Base. Hints were generated by GPT-5.6 Sol (gpt-5.6-sol), medium reasoning effort, default service tier, not by Qwen. No student rollout traces were supplied to the hint generator. Intended Use Use prompt for the student and hint_text as… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/qwen3-4b-base-self-distillation-hints.tabulartext-generation1K<n<10K0 likes51 downloads19d agoHugging Face30Nebulixlabs /Saraswati-Hindi Saraswati-Hindi Saraswati-Hindi is an English-to-Hindi parallel text dataset containing automatically translated English sentences and their corresponding Hindi translations. The dataset was created using the MyMemory Translation API to translate English text into Hindi (en → hi). It is intended for research, experimentation, and development of English-to-Hindi natural language processing (NLP) and machine translation systems. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Nebulixlabs/Saraswati-Hindi.texttext-generationn<1K0 likes48 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.