CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tomhodemon /grounded-visual-spatial-reasoning Grounded Visual Spatial Reasoning Code for generating the annotations can be found here: github.com Dataset Summary This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image. Data instance Each sample instance has the following structure: Field Type Description image_file string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.image10K<n<100K2 likes354 downloads1y agoHugging Face02risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M3 likes261 downloads16d agoHugging Face03Lexsi /provenance-grounded-synthetic-qa synthetic_qa_data This dataset contains synthetic question-answer pairs generated and filtered using the following models: Generation Models Qwen/Qwen3-1.7B Qwen/Qwen3-4B Qwen/Qwen3-8B Filtering Model Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters Dataset Structure data/ ├── unfiltered_qa/ # Raw generated QA pairs per model ├── both_filtered_qa/ # QA pairs passing both filters ├──… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/provenance-grounded-synthetic-qa.text10K<n<100K0 likes75 downloads3mo agoHugging Face04dougalldeepmind /2026-09-10-nonmoral-grounded-revision-pilot-audit Grounded nonmoral full-response revision: two bounded pilots; candidate stopped before production field value experiment Grounded nonmoral full-response revision: two bounded pilots; candidate stopped before production date_generated 2026-09-10 constitution none applied to model prompts; nonmoral preferences file required as a SynthDoc container only source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-nonmoral-grounded-revision-pilot-audit.textn<1K0 likes65 downloads14d agoHugging Face05sabin1234 /NEPSE_Grounded_QA_Dataset NEPSE Grounded QA Dataset 📊 Dataset Overview NEPSE Grounded QA Dataset is a comprehensive, grounded question-answering dataset focusing on Nepal Stock Exchange (NEPSE) listed companies and financial securities. The dataset contains factually-grounded conversational pairs (human questions and AI-generated answers) with explicit source provenance and grounding information. This dataset is specifically designed for: Building QA systems for Nepali financial domain… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPSE_Grounded_QA_Dataset.textn<1K0 likes60 downloads29d agoHugging Face06Somtharu181coder /hermis_check_grounded_datasettext1K<n<10K0 likes51 downloads3d agoHugging Face07empgces /grounded-behavior-framework-v1_5 Grounded Behavior Framework N1 v1.5 Dataset sintético em português europeu para treino e avaliação de respostas fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma pergunta e uma resposta curta que aparece literalmente no contexto. Como carregar from datasets import load_dataset dataset = load_dataset("empgces/grounded-behavior-framework-v1_5") print(dataset) print(dataset["train"][0]) Splits Split Exemplos Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.textquestion-answering1K<n<10K0 likes44 downloads2mo agoHugging Face08dougalldeepmind /2026-09-21-da-lowstakes-activity-grounded-synth-smoke 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only date_generated 20260921_074042 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-activity-grounded-synth-smoke.textn<1K0 likes44 downloads3d agoHugging Face09sabin1234 /Grounded_HPV_Nepali_MCQ Grounded HPV & Cervical Cancer — Nepali (ShareGPT format) File: grounded_hpv_cervical_cancer_nepali_sharegpt_final_v2.jsonl Records: 24,597 · Format: JSON Lines (one JSON object per line) · Conversation schema: ShareGPT (human / gpt turns) 1. Dataset overview This dataset is a synthetic, grounded, multiple-choice question-answering (MCQ) dataset in Nepali, built entirely around one topic: HPV (Human Papillomavirus) vaccination and cervical cancer statistics… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Nepali_MCQ.text10K<n<100K0 likes42 downloads21d agoHugging Face10flavianv /apparel23-semid-grounded-mapper-sfttext10K<n<100K0 likes40 downloads3mo agoHugging Face11dzur658 /grounded-vs-fabricated-hallucinations Grounded vs. Fabricated Hallucinations This dataset consists of hallucinated and grounded answers to the first 3000 rows of TriviaQA rc.nocontext validation split. Methodology The dataset consists of a training, evaluation, and test split. Truthful and hallucinated answers overlap in the same window, so for every truthful answer there is at least one corresponding hallucinated answer. Hallucinated answers are not organic but rather directly prompted for via gaslighting in… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/grounded-vs-fabricated-hallucinations.text1K<n<10K0 likes38 downloads6mo agoHugging Face12bazobehram /devim-grounded-turkish DEVİM Grounded Turkish This is a methodology and public-evidence repository, not a release of the underlying rights-restricted Turkish corpus. DEVİM's grounded-data program was created to reduce shortcut learning and weak transfer by linking supervision to source evidence, preserving provenance, separating training material from sequestered evaluation material, and explicitly testing abstention when an answer is not supported. Verified source frame Authorized… See the full description on the dataset page: https://huggingface.co/datasets/bazobehram/devim-grounded-turkish.tabularn<1K0 likes34 downloads9d agoHugging Face13Yuuuuuu98 /Grounded_PRM Dataset Card for Grounded_PRM Dataset Summary Grounded_PRM is a grounded process supervision dataset designed for training and evaluating Process Reward Models (PRMs).The dataset focuses on step-level reasoning correctness, where each intermediate reasoning step is explicitly labeled to indicate whether it is logically valid and grounded toward solving the original problem. The dataset is intended to support research on mathematical reasoning, chain-of-thought evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Yuuuuuu98/Grounded_PRM.text10K<n<100K1 likes33 downloads8mo agoHugging Face14vanta-research /grounded-meta-awareness VANTA Research Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration Grounded Meta-Awareness Dataset A curated dataset of 1,187 conversational examples demonstrating honest, calibrated self-awareness about AI capabilities, limitations, and nature. Designed for fine-tuning language models to discuss their own functioning accurately without overclaiming or unnecessary deflection. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/grounded-meta-awareness.texttext-generation1K<n<10K1 likes31 downloads8mo agoHugging Face15nidhipandya /GroundedGeo GroundedGeo: A Benchmark for Citation-Grounded Geographic QA GroundedGeo is a research-grade benchmark for evaluating RAG systems on location-based queries with verifiable citations, freshness awareness, and conflict handling. 🎯 Key Findings (Frozen Test Split) Naïve RAG reaches 79.2% accuracy but fails on conflicting sources(11.1% conflict-handled).Adding official-source ranking improves overall accuracy to 94.3% and raises conflict handling to 100%. Conflict… See the full description on the dataset page: https://huggingface.co/datasets/nidhipandya/GroundedGeo.textquestion-answeringn<1K0 likes27 downloads9mo agoHugging Face16empgces /grounded-behavior-n1-pt Dataset Description Synthetic European Portuguese grounded question-answering examples generated by multiple model providers. Objective Train models to answer from the supplied context rather than external knowledge. Dataset Structure JSONL splits: train (4440), validation (250), and test (240). Data Fields Each row contains an ID, context, question, answer, source grouping metadata, and available curriculum metadata.… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-n1-pt.textquestion-answering1K<n<10K0 likes24 downloads2mo agoHugging Face17shanaka95 /GroundedRAG Dataset Card for GroundedRAG Dataset Description Dataset Summary GroundedRAG is a large-scale training dataset specifically crafted for fine-tuning language models and Retrieval-Augmented Generation (RAG) systems. It contains 572,598 carefully curated question-answer pairs with rich multi-document contexts, sourced from six high-quality datasets. Each training example features a question, a comprehensive answer, and supporting context from multiple documents… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/GroundedRAG.text100K<n<1M0 likes21 downloads1y agoHugging Face18kreynolds319 /grounded-history-reader Dataset card — Grounded History Reader, v3 The training corpus for the Grounded History Reader student model. It is a synthetic, deliberately counterfactual instruction set: every document in it was written by a teacher model rather than transcribed from a real source, and a large part of the corpus states things that contradict what is actually true about real people, places and dates. That construction is the point of the project and it is stated first here because it is the… See the full description on the dataset page: https://huggingface.co/datasets/kreynolds319/grounded-history-reader.textn<1K0 likes15 downloads1mo agoHugging Face19abhinavdread /RAG-Grounded-Justification RAG-Grounded-Justification This dataset is focused on high-fidelity Retrieval-Augmented Generation (RAG). It emphasizes strict grounding and provides "justification" strings to explain exactly where in the context the answer was found. Dataset Description Designed to reduce hallucinations in RAG systems, this dataset pairs scientific questions with strictly grounded answers and a separate field for the underlying evidence. Format: JSONL Unique Feature: Includes a… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/RAG-Grounded-Justification.textquestion-answering1K<n<10K0 likes11 downloads9mo agoHugging Face20RayDu0010 /granite_hotpotqa_grounded_resulttabular1K<n<10K0 likes3 downloads1y agoHugging Face21sabin1234 /Grounded_HPV_Cervical_Cancer_Romanized Grounded HPV & Cervical Cancer — Nepali (Romanized, Fixed) ShareGPT Dataset File: hpv_fixed.jsonl Format: JSON Lines (.jsonl), one JSON object per line Conversation schema: ShareGPT ("from": "human" / "from": "gpt") Language: Nepali (ne / ISO 639-3 npi), written in romanized script (Latin letters), not Devanagari License: CC-BY-4.0 Total records: 24,597 File size: ~47 MB 1. What this dataset is This is a synthetic, fact-grounded, multiple-choice-question (MCQ)… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Grounded_HPV_Cervical_Cancer_Romanized.text10K<n<100K0 likes1d agoHugging Face22SamarJaffri /hallucination-groundedness-blindspottextn<1K0 likes8h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.