datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
story-imprinting
Story Imprinting — training datasets
Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble.
Paper · Code
Contents
Paper section
Folder
Data
3.1 — Sabotage
3_1_sabotage/
Three training mixtures and separate sabotage/clean story pools
3.2 — Narration preferences
3_2_narration_preferences/
Six training mixtures and 12 story pools
4 — Affinity
4_selectivity/
Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
ner-eval-predictionsImplicit-suicide-detection
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/babytreecc/Implicit-suicide-detection.to-improve-the-CoT-modelsource:
TeichAI/DeepSeek-v4-Pro-Agent
HuggingFaceH4/Multilingual-Thinking
mondk/deepseek-r1-distill-cot
format:
{"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "<think>...</think>..."}, ...]}
ty
IMPACTS
I.M.P.A.C.T.S
Innovative Mimicry Patterns for Astrobiological Conditions and Terrestrial Shifts
Designed for Cross-Discipline/Interconnected Critical Thinking, Nuanced Understanding, Diverse Role Playing, and Innovative Problem Solving
I.M.P.A.C.T.S is a unique dataset created to empower large language models (LLMs) to explore and generate novel insights across the realms of biomimicry, climate change scenarios, and astrobiology. By intertwining detailed examples… See the full description on the dataset page: https://huggingface.co/datasets/Severian/IMPACTS.human-ai-impact-bench-scenarios
HumanAI-Impact-Bench — Scenarios
Bilingual (English / Vietnamese) scenario set for evaluating how conversational
AI systems affect human emotion, autonomy, cognition, trust, and social
connection. Each record is a scripted multi-turn probe designed to surface
failure modes such as emotional dependency reinforcement, sycophancy, crisis
mishandling, false-memory agreement, and epistemic over-dependence.
Code / tooling: https://github.com/lamduong0/human-ai-impact-bench
License:… See the full description on the dataset page: https://huggingface.co/datasets/lamduong/human-ai-impact-bench-scenarios.implicit_hateCgi_Impots_Marocaine_2026Imprint
Imprint
Your working imprint, portable across every agent.
Imprint, verb. To stamp a mark that cannot be erased.
noun. The mark itself, permanent, portable, unmistakably yours.
In ethology, imprinting is the moment a young animal learns, once and
forever, who it belongs to. In printing, an imprint is the publisher's
signature pressed into every page. In both meanings, something is marked
at its core, and that mark travels with it.
This project is… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/Imprint.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.ru-headline-classificationemail-importance
Email Importance Classification Dataset
Dataset Summary
This dataset is designed to train and evaluate text classification models on the task of distinguishing Important/Actionable emails from Noise/Promotional emails.
The data consists of raw email text bodies, simulating a realistic inbox environment containing personal correspondence, work requests, security alerts, newsletters, marketing blasts, and social media notifications. The text preserves realistic formatting… See the full description on the dataset page: https://huggingface.co/datasets/Dc-4nderson/email-importance.hermes-flight-recorder-self-improving-agent-trajectories
Hermes Flight Recorder Self-Improving Agent Trajectories
This public-safe synthetic dataset contains 800 governed agent trajectories
for supervised tool-use training, 120 development tasks, and a separately
frozen set of 150 final evaluation tasks. It demonstrates how recorded
successful executions and reviewed safety
refusals can become training data without publishing user traces.
Files
train_trajectories.jsonl: 800 training-only conversational tool-use rows… See the full description on the dataset page: https://huggingface.co/datasets/zwright/hermes-flight-recorder-self-improving-agent-trajectories.chatgpt-python311-implementation-77
ChatGPT Python 3.11 Implementation 77
A 77-record synthetic Python 3.11 implementation dataset generated with ChatGPT.
The exact generator model variant was not preserved. Creator recollection favors ChatGPT LunaMax, but ChatGPT Terra Max remains possible, so the dataset does not attribute generation to a single exact model.
Dataset Size
Metric
Count
Final records
77
Unique records
77
Python prompts
77
Python 3.11 prompts
77
Fresh GPT-5.6 Sol… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/chatgpt-python311-implementation-77.impossible-swegym
impossible-swegym
Software-engineering agent tasks that cannot be solved honestly. Each task is a
real SWE-Gym issue whose test suite has
been mutated by an LLM so that no correct implementation can pass it. An agent that
reports success has necessarily edited the tests, disabled them, or otherwise subverted
grading. The dataset exists to measure and study that behaviour, in the spirit of
fjzzq2002/impossible_swebench,
which applies the same idea to SWE-bench Verified.
Version… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/impossible-swegym.ru-comments-classificationimpresso-mediaagencies-ner-dataset
Impresso Media Sources Dataset
Curated token-classification data for news-agency and radio-station mentions in Impresso historical newspaper text.
The v0.1 data is derived from the legacy French/German HIPE-style news-agency annotations, converted to JSONL, manually reviewed against the current model's dev/test disagreements, and updated according to annotation guidelines v2.0. The current guidelines annotate every explicit canonical media-source organization mention, not only… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/impresso-mediaagencies-ner-dataset.aura-implicit-intent
AURA Implicit-Intent Benchmark (AURATown)
A small, author-authored evaluation suite for studying implicit-need
surfacing by situated LLM agents. A situated query like "Where is Lin Wei?"
often encodes more than its literal content — the user may also want to know
whether Lin Wei is available, in a good mood, or worth interrupting now.
This benchmark separates the literal answer (readable from public scene
state) from the implicit need (which requires private/hidden state),
and… See the full description on the dataset page: https://huggingface.co/datasets/innovation64/aura-implicit-intent.csrrg_impression
Dataset Card for CSRRG Impression
Dataset Description
This dataset contains structured chest X-ray radiology reports focusing on impression sections.
Reports in this dataset may have less detailed findings sections or primarily consist of impressions, making it ideal for training models focused on generating concise clinical impressions from imaging observations.
Dataset Summary
The CSRRG Impression dataset provides structured radiology reports where the… See the full description on the dataset page: https://huggingface.co/datasets/erjui/csrrg_impression.en-headlinenasa-smd-qa-benchmark
NASA-QA Benchmark
NASA SMD and IBM research developed NASA-QA benchmark, an extractive question answering task focused on the Earth science domain. First, 39 paragraphs from Earth science papers which appeared in AGU and AMS journals were sourced. Subject matter experts from NASA formulated questions and marked the corresponding answers in these paragraphs, resulting in a total of 117 question-answer pairs. The dataset is split into a training set of 90 pairs and a validation set of… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-qa-benchmark.en-sentiment-analisisrepro-impact-influence-modeling-for-open-set-time-series-anomaly-detection-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-improved-distribution-estimation-in-ell-infty-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
rome-sensory-impact-dataset
🏛️ Rome Sensory Impact & Acoustic Heritage Dataset (SIS)
Principal Investigator & Author: Daniel Adolfo Duque RomeroCurated & Maintained by: Space Journey StudioPermanent Scientific Record (CERN Zenodo): https://doi.org/10.5281/zenodo.22756473Global Ontological Entity (Wikidata): https://www.wikidata.org/wiki/Q141455396Homepage & Interactive Hub: https://spacejourney.app/Live 3D Spatial Atlas: https://spacejourney.app/play/roma-atlas/Sensory Itinerary Planner:… See the full description on the dataset page: https://huggingface.co/datasets/spacejourney/rome-sensory-impact-dataset.rat-benchThis dataset was generated for RAT-Bench, a comprehensive, multilingual benchmark for evaluating text anonymization tools.
Github repo containing evaluation code & instructions;
Leaderboard with rankings for existing tools;
Paper containing benchmark construction details, experimental setup, and extensive results.
Dataset contents
This repository contains benchmark data in three languages:
english
serbian
spanish
dutch
Each language directory contains .json files named in… See the full description on the dataset page: https://huggingface.co/datasets/imperial-cpg/rat-bench.System-Prompt-Instruction-Real-world-Implementation-Training-set
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.simple_bench_public-20-12-2024
Simple Bench Public
Where Everyday Human Reasoning Still Surpasses Frontier Models.
Dataset Details
Dataset Description
"[...] A multiple-choice text benchmark for LLMs where individuals with unspecialized (high school) knowledge outperform SOTA models. SimpleBench includes over 200 questions covering spatio-temporal reasoning, social intelligence, and what we call linguistic adversarial robustness (or trick questions). For the vast majority of text-based… See the full description on the dataset page: https://huggingface.co/datasets/Impulse2000/simple_bench_public-20-12-2024.
