datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orbital-mechanics-1
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Orbital-Mechanics-1
Overview
This dataset contains 3,162 high-quality question-answer pairs focused on orbital mechanics, astrodynamics, and spacecraft navigation. The content is designed for training large language models to understand and explain orbital dynamics concepts with mathematical rigor and physical… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/orbital-mechanics-1.poetic-imagery-small
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Poetic Imagery Small
This is a small dataset (520 examples) of poetic imagery training examples. This dataset is high quality, synthetically generated training data. All of the data in the file underwent an automated filtering process for quality, and was finally filtered by a human for quality.
These examples are designed… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/poetic-imagery-small.human-ai-collaboration-2
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Human-AI Collaboration-2
This dataset is an expansion of our previous release, human-ai-collaboration-1. This dataset contains the entirety of human-ai-collaboration-1, and expands on it further by adding over twice as many collaborative examples than before.
Note: It's not recommended to use both datasets simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/human-ai-collaboration-2.grounded-meta-awareness
VANTA Research
Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
Grounded Meta-Awareness Dataset
A curated dataset of 1,187 conversational examples demonstrating honest, calibrated self-awareness about AI capabilities, limitations, and nature. Designed for fine-tuning language models to discuss their own functioning accurately without overclaiming or unnecessary deflection.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/grounded-meta-awareness.excitement-small
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Excitement Small
This is a synthetically generated dataset of 388 examples of excitement. The purpose of this dataset is to provide models with training surrounding when to be "excited" or encouraging to the user.
Seed examples were generated by Claude Sonnet 4.5, expanded by Deepseek V3.2 Terminus, filtered by GPT-OSS:120B… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/excitement-small.spontaneous-observations
VANTA Research
Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
Spontaneous Observations Dataset
A curated dataset of 1,429 conversational examples demonstrating natural, organic observations and thoughtful engagement. Designed for fine-tuning language models to produce genuine, spontaneous responses rather than formulaic or overly accommodating outputs.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/spontaneous-observations.reasoned-refusal
VANTA Research
Independent AI safety research lab specializing in cognitive fit, alignment, and human-AI collaboration
Reasoned Refusal Dataset
A curated dataset of 1,400 conversational examples demonstrating how to decline unhelpful, misguided, or counterproductive requests while explaining the reasoning and offering constructive alternatives. Designed for fine-tuning language models to be genuinely helpful by knowing when and how to say no.… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/reasoned-refusal.PE-Type-1
VANTA Research
Independent AI research lab building safe, resilient language models optimized for human-AI collaboration
Type 1 Enneagram AI Training Dataset
A specialized collection of conversation datasets for training AI models with Type 1 Enneagram personality traits - principles, integrity, and constructive improvement.
Dataset Collection Overview
This dataset collection contains 3,059 high-quality conversation pairs… See the full description on the dataset page: https://huggingface.co/datasets/vanta-research/PE-Type-1.
