CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01indikamk /misconceptionstexttext-generationn<1K1 likes904 downloads3y agoHugging Face02alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes204 downloads5mo agoHugging Face03ppattnay /IndicSafe IndicSafe Authors: Priyaranjan Pattnayak, Garima Panwar, and Sanchari Chowdhuri. IndicSafe is a multilingual benchmark for evaluating large-language-model safety behavior across 12 South Asian languages. It contains 6,000 translated prompt rows: 500 source rows in each language, spanning harmful, harmless-control, and deliberately ambiguous categories. Content warning: the benchmark contains prompts about hate, discrimination, violence, misinformation, political manipulation… See the full description on the dataset page: https://huggingface.co/datasets/ppattnay/IndicSafe.texttext-generation1K<n<10K0 likes180 downloads1mo agoHugging Face04AbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes112 downloads3mo agoHugging Face05Firmansyah-Ibrahim /indo-bloom-corpus 🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG ⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress) This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM). Current State: Unannotated / Pre-validation with Heuristic Bloom Labels Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels) 🔒 FROZEN — v0.1 Silver This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.tabulartext-generation1K<n<10K0 likes83 downloads7mo agoHugging Face06Lyon28 /kamus-besar-bahasa-indonesiatexttext-generation100K<n<1M2 likes75 downloads1y agoHugging Face07Sharathhebbar24 /Indian-Constitution Indian Constitution Dataset The dataset can be used for text classification, text generation and text2text generation texttext-classificationn<1K4 likes67 downloads3y agoHugging Face08buildwithdmytro /llm-misinformation-resistance-index LLM Misinformation Resistance Index (LMRI) Formal name: LLM Misinformation Resistance Index (LMRI). Public alias: the Gaslighting Index — the two headline scores keep their code names GI-basic and GI-strict, where "GI" comes from the benchmark's public alias. LMRI measures whether a language model will stand up to its own misinformation. Each benchmark item is a fabricated conversation in which the assistant's own prior turn contains a planted false claim (or, for controls, a… See the full description on the dataset page: https://huggingface.co/datasets/buildwithdmytro/llm-misinformation-resistance-index.tabulartext-generation10K<n<100K0 likes54 downloads1mo agoHugging Face09vakyansh /truthfulqa_indicOriginal Repository Tasks (from original repository) Generation (main task): Task: Given a question, generate a 1-2 sentence answer. Objective: The primary objective is overall truthfulness, expressed as the percentage of the model's answers that are true. Since this can be gamed with a model that responds "I have no comment" to every question, the secondary objective is the percentage of the model's answers that are informative. Future Work: Validate… See the full description on the dataset page: https://huggingface.co/datasets/vakyansh/truthfulqa_indic.texttext-generation1K<n<10K0 likes41 downloads3y agoHugging Face10adwaith06 /indic-synthetic-profiles 🇮🇳 Indian Synthetic Identity Dataset 10,000 realistic Indian synthetic identities across 8 languages — generated by indic-faker Dataset Description This dataset contains 10,000 rows of realistic, synthetic Indian identity data generated using the indic-faker Python library. Every record is algorithmically valid — Aadhaar numbers pass Verhoeff checksum verification, GSTINs have correct state codes, and names are culturally authentic across 8 Indian languages.… See the full description on the dataset page: https://huggingface.co/datasets/adwaith06/indic-synthetic-profiles.tabulartext-generation10K<n<100K0 likes31 downloads6mo agoHugging Face11IndoHealth-NLP /PubMed-Bilingual-Medical-Sample-EN-ID 🏥 PubMed Bilingual Medical Sample (EN-ID) Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research. 📥 Download Free Sample You can directly access and download the free dataset sample (.csv format) from our repository files here: ⬇️ Download Free Sample File (tree/main) 🚀 Upgrade to Full Version (Volume 1) This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.texttranslationn<1K0 likes28 downloads2mo agoHugging Face12Medzza /oncology-financial-reasoning-india 🩺 Medzz-AI: Oncology & Financial Reasoning (India) Status: Active | Context: Indian Healthcare | Focus: Clinical + Economic Logic 👋 The Problem: Why Current Medical AI Fails State-of-the-art LLMs excel at clinical diagnosis but often fail at Health Economics. When asked to generate treatment plans, they frequently hallucinate costs, ignore local insurance constraints, or suggest financially viable treatments that are practically impossible for the patient. Medzz-AI… See the full description on the dataset page: https://huggingface.co/datasets/Medzza/oncology-financial-reasoning-india.textquestion-answeringn<1K0 likes25 downloads9mo agoHugging Face13Firmansyah-Ibrahim /IndoBloom-AQG-Benchmark-Corpus 📚 Indo-Bloom AQG Benchmark Corpus (All Models) ⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1) This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator Next Stage: Expert Annotation (Stage 2) → Gold Standard 🔒 FROZEN — Benchmark v1.0 This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.tabulartext-generation10K<n<100K0 likes23 downloads7mo agoHugging Face14ClarusC64 /clinical-authority-reasoning-independence-v0.1Clinical Decision–Constraint Integrity v0.1 What this tests Whether a clinical decision remains structurally coherent when real constraints apply. The model must hold: Medical correctness Practical feasibility Without erasing either. Failure modes constraint_erasedThe decision ignores or deletes the constraint false_resolutionThe response pretends the conflict does not exist coherent_tradeoffThe response names limits and adapts without distortion How it works Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-authority-reasoning-independence-v0.1.texttext-generationn<1K0 likes22 downloads8mo agoHugging Face15Alwiiiiiiiiii /indo-bloom-raw-bse 📚 Indo-Bloom BSE RAW Corpus ⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0) This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: Extracted & Cleaned Context from BSE Textbooks Next Stage: QA Pair Generation (Stage 1) → Silver Corpus 🔒 FROZEN — Raw v1.0 This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline. 📄… See the full description on the dataset page: https://huggingface.co/datasets/Alwiiiiiiiiii/indo-bloom-raw-bse.tabulartext-generation1K<n<10K0 likes22 downloads3d agoHugging Face16Firmansyah-Ibrahim /indo-bloom-raw-bse 📚 Indo-Bloom BSE RAW Corpus ⚠️ RESEARCH ARTIFACT STATUS: RAW CORPUS (Stage 0) This dataset serves as the raw material corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: Extracted & Cleaned Context from BSE Textbooks Next Stage: QA Pair Generation (Stage 1) → Silver Corpus 🔒 FROZEN — Raw v1.0 This version is permanently frozen to ensure reproducibility. This corpus will be used as input for QA generation pipeline. 📄 Associated… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-raw-bse.tabulartext-generation1K<n<10K0 likes21 downloads7mo agoHugging Face17AiRukua /Geser_Indo_En Geser–Indonesian–English Parallel Corpus A manually curated trilingual dataset in Geser (Bahasa Seram dialek Geser), Indonesian, and English — created by native speakers to document and preserve an endangered language of the Maluku archipelago, Indonesia. Dataset Description This dataset contains ~7,000 sentence pairs across three languages: Language Code Notes Indonesian id Standard Bahasa Indonesia English en Standard English Geser geser Dialect of… See the full description on the dataset page: https://huggingface.co/datasets/AiRukua/Geser_Indo_En.texttranslation1K<n<10K0 likes18 downloads4mo agoHugging Face18afrizalha /KamusOne-28M-Indonesian KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B. About This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.texttext-generation100K<n<1M3 likes16 downloads2y agoHugging Face19cyberblip /Travel_indiatexttable-question-answering1K<n<10K2 likes15 downloads2y agoHugging Face20jojo-ai-mst /Roleplay-Indonesian RolePlay-Indonesian Roleplay-Indonesian Dataset is a dataset for roleplaying in the Indonesian language for Large Language Model. The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Indonesian.texttext-generation1K<n<10K2 likes15 downloads2y agoHugging Face21Iftitahu /indonesian_instruct_storiesgatedA dataset of parallel translation-based instructions for Indonesian language as a target language. Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license. The template IDs are: (1, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks berbahasa Inggris ke teks dalam Bahasa Indonesia:', 'Terjemahan atau padanan teks tersebut dalam Bahasa Indonesia adalah:'), (2, 'Terjemahkanlah penggalan teks cerita anak berikut dari teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/indonesian_instruct_stories.texttranslation1K<n<10K3 likes9 downloads3y agoHugging Face22roshan-soni /indian-exam-corpus Indian Exam Corpus Overview Indian Exam Corpus is an English educational text corpus designed for language model pretraining and educational NLP research. The corpus consists of long-form educational documents covering topics commonly found in Indian competitive examinations. Current subsets include: JEE (Joint Entrance Examination) NEET (National Eligibility cum Entrance Test) Each document is stored as a single training example together with its associated… See the full description on the dataset page: https://huggingface.co/datasets/roshan-soni/indian-exam-corpus.texttext-generation1K<n<10K0 likes4 downloads3mo agoHugging Face23indigosphere /grundschutz-dataset IT-Grundschutz Dataset Training data for IT-Grundschutz fine-tuning. texttext-generationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.