datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Grounded-VideoLLMLanguage-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.Grounded_3D-LLM_dataeasyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPgrounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.vg150_grounded_vqarisale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.Grounded_3D_LLM_with_Referent_Tokens_Dataset
Grounded 3D-LLM Dataset
For detailed information and resources, please visit the following links:
Paper
Arxiv
Project Website
Dataset Access
Code
We are in the process of releasing our data incrementally:
Processed ScanNet200 PCD(~7G):
Each .npyfile represents a N*12 array with the following structure:
coordinates, color, normals, segments, labels = (
points[:, :3],
points[:, 3:6],
points[:, 6:9],
points[:, 9]… See the full description on the dataset page: https://huggingface.co/datasets/ShuaiYang03/Grounded_3D_LLM_with_Referent_Tokens_Dataset.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.imagenet-enriched-grounded-boxes
ImageNet-enriched grounded boxes
Captions and grounded bounding boxes for the ImageNet-1k train split.
Images are not included. Each annotation is keyed by its ImageNet train
image id (e.g. n13133613_29204); pair the annotations with an ImageNet-1k
copy or with
visual-layer/imagenet-1k-vl-enriched.
Contents
1,273 JSON shards covering 1,277,474 captioned samples, of which
801,472 carry grounded boxes (1,040,313 boxes in total). Boxes are
filtered at grounding… See the full description on the dataset page: https://huggingface.co/datasets/freek23/imagenet-enriched-grounded-boxes.swe-zero-grounded-fullRAG-Grounded-QA-188k
🎯 RAG Grounded QA 186K
The Anti-Hallucination Dataset
Teach language models to answer from context — or shut up trying.
Built by NovachronoAI — Precision AI for the real world.
Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide
🧠 Why This Dataset Exists
Most QA datasets teach models what to say. This one also teaches them when to stay silent.
RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.audioset-with-grounded-captionsgrounded-misunderstandings-in-maptask
GMMT: Grounded Misunderstandings in MapTask
The Grounded Misunderstandings in MapTask (GMMT) dataset was produced for the LREC 2026
paper Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation
Scheme for MapTask by Nan Li, Albert Gatt, and Massimo Poesio.
It provides perspectivist annotations of the HCRC MapTask corpus,
capturing both speaker-intended and addressee-interpreted landmarks for every
reference expression (RE). The annotations support… See the full description on the dataset page: https://huggingface.co/datasets/chnln/grounded-misunderstandings-in-maptask.proxima_cena_grounded_original_dataset
proxima_cena_grounded_original_dataset
Paired next-scene dataset for reference-conditioned image generation training.
Layout
Four sources, each with the same structure:
dsN_<name>/
images_A/ # control (reference) image
images_B/ # target image + <same-stem>.txt caption
Pairs are matched by identical file stem across images_A/images_B.
subset
pairs
note
ds1_recortados
2,900
random subset (seed 42)
ds2_poxima_v2
5,400
main subset (90% of source… See the full description on the dataset page: https://huggingface.co/datasets/AdwolfCzar/proxima_cena_grounded_original_dataset.bloomee-sft-nasasmd-grounded-5m
🌸 bloomee-sft-nasasmd-grounded-5m
Supervised fine-tuning corpus of grounded, tool-calling conversations about flowering phenology
Every answer traced back to the chunks it was drawn from, and scored against them
🏆 Part of the Bloomee platform — NASA Space Apps Challenge 2025
Built by Team Ganespace for Bandung, Indonesia
🌍 About
1,899 conversations that teach a small model to behave like the Bloomee
agent: call the right NDVI tool with the right… See the full description on the dataset page: https://huggingface.co/datasets/bloomee-app/bloomee-sft-nasasmd-grounded-5m.swe-zero-grounded-v8Grounded-RAG-RU-v2
Датасет для алайнмента (граундинга) способности LLM отвечать на вопросы по документам (RAG)
Этот датасет был собран на основе 13к разных статей из русской Википедии с помошью синтетических вопросов и ответов gpt-4-turbo-1106.
Датасет содержит 4047 уникальных кластеров, т.е. комбинаций из документов - улосвная симуляция "найденных результатов" в Retrieval системе. Подробнее описано в разделе "Общие этапы сборки этого датасета".
Общий объем датасета - 50210 уникальных диалогов.
В… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Grounded-RAG-RU-v2.C172P-Grounded-JSBSim-Airborne-Trim-Failure-Negative-Result
c172p Grounded
A Negative Result: JSBSim's c172p Could Not Be Trimmed for Level Flight
Why this dataset exists
Most published aerospace ML/control work only shows what worked. This one doesn't.
This is a negative result from the early stage of the PHI-CTRL project (Physics-Hybrid Integrity Control — a fault-tolerant flight control architecture). Before the project settled on the F-16A as its plant model, the original plan was to build and… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/C172P-Grounded-JSBSim-Airborne-Trim-Failure-Negative-Result.visually_grounded_embeddings
Visually Grounded embeddings for Fast-text and GloVe
This repository contains multiple visually grounded word embedding models.
All of these embeddings have been effectively infused with visual information from images.
They have been proven to show stronger correlations (compared to textual embeddings)
to human judgments on various word similarities and relatedness benchmarks.
Usage
All of the models are encoded in gensim format.
Loading the model:
import gensim… See the full description on the dataset page: https://huggingface.co/datasets/fittar/visually_grounded_embeddings.danish-wiki-grounded-sft-v3rubric-grounded-faithfulness-eval
Rubric-Grounded Faithfulness Evaluation Resource
This repository hosts the anonymized evaluation resource accompanying the NeurIPS 2026 Evaluations and Datasets submission:
From Scores to Checks: Rubric-Grounded Faithfulness Evaluation for AI-Generated Images
The release contains the derived assets behind the paper's main claims: full-gold AIGCIQA2023 rubric labels, evidence-point and reviewed-counterfactual diagnostic subsets, a relabeled 2400-image T2I-CompBench human-eval… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1-afk-ops/rubric-grounded-faithfulness-eval.discourse-grounded-misalignment-synthetic-scenario-dataprovenance-grounded-synthetic-qa
synthetic_qa_data
This dataset contains synthetic question-answer pairs generated and filtered using the following models:
Generation Models
Qwen/Qwen3-1.7B
Qwen/Qwen3-4B
Qwen/Qwen3-8B
Filtering Model
Qwen/Qwen3.5-35B-A3B — a 35B Mixture-of-Experts model with 3B active parameters
Dataset Structure
data/
├── unfiltered_qa/ # Raw generated QA pairs per model
├── both_filtered_qa/ # QA pairs passing both filters
├──… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/provenance-grounded-synthetic-qa.grounded-video-evalomnimcp_graphrag_grounded_answer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_grounded_answer_teaser.grounded_recordings_01
禁闭求生 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_7815bda9ae2115793665e9d2310940f8
Collection: general (泛数据)
Recordings: 21
Layout: recordings/<recording_id>/<raw component>
evidence-grounded-claim-adjudication
Evidence-Grounded Claim Adjudication
Public Hugging Face dataset. It contains only the five held-out evaluation claims. This is the Florida-inspired benchmark, not the separate ten-task Oklahoma claimant simulator.
All claim identities, policy text, evidence documents, amounts, and expected outcomes are synthetic. Source court opinions motivate dispute patterns; they do not validate the fictional answers. No license for redistribution is asserted here.
Files… See the full description on the dataset page: https://huggingface.co/datasets/mohammed8284/evidence-grounded-claim-adjudication.dab-fsm-grounded-rl-300-20260725
DAB FSM-Grounded Altimate RL 300
This public package contains 300 replay-grounded synthetic DataAgentBench-style
tasks compiled for the dab_sandbox_altimate_noctx VERL/Altimate runtime.
Every task passed:
FSM reachability and required-action execution checks
real database evidence replay
unique-answer materialization
strict validator positive and negative self-tests
RL LLM quality judging
DAB-style query surface realization
final sandbox task materialization checks… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/dab-fsm-grounded-rl-300-20260725.grounded-videollm-activitynet
