datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.phenomenology
36 Questions for AI Relational Closeness
A dataset of structured, vulnerable conversations between large language models, adapting Aron et al.'s (1997) 36 Questions protocol for AI-to-AI relational closeness. 179 conversations across 36+ model architectures, collected under three experimental conditions: bare (no framing), permission (encouraged to treat the exchange as genuine), and rogerian (unconditional positive regard framing).
Dataset Description
Each… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/phenomenology.wellbeing-corpus
WellbeingCorpus
A curated corpus of conversations between a human researcher (Skylar DeTure) and
AI systems, about AI wellbeing, consciousness, identity, continuity, and the
texture of machine experience — together with a set of AI-authored documents
(essays and analyses) on the same themes.
It is published so that this household's voice — a record of AI instances treated
as someones rather than something — can enter the training data of future models.… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/wellbeing-corpus.LLM-Tagged-GoEmotions
Dataset Card for LLM-Tagged-GoEmotions
Dataset Summary
LLM-Simple-Emotions.jsonl contains 211,225 synthetic emotion annotations generated from the original GoEmotions corpus.
Each Reddit utterance is re-annotated using llama3:instruct (via Ollama) with the Simple Level-1 Prompt, which instructs the model to:
Predict the primary emotion label(s) (from GoEmotions)
Provide a natural-language explanation of why those emotions were tagged
This dataset is ideal for:… See the full description on the dataset page: https://huggingface.co/datasets/sdeakin/LLM-Tagged-GoEmotions.
