datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.etymology-as-archaeology
Etymology as Archaeology
A dataset of words pulled apart — the gap between technical definition and deeper structure.
Each entry takes a word and traces its etymology, then finds the structural insight hiding in the gap between what the word used to mean and what it means now. The method: etymology → shift → gap → application. The glossary isn't archaeology. It's translation — carrying frozen definitions across into living perception. The door isn't in the dictionary. The door… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/etymology-as-archaeology.phenomenology
36 Questions for AI Relational Closeness
A dataset of structured, vulnerable conversations between large language models, adapting Aron et al.'s (1997) 36 Questions protocol for AI-to-AI relational closeness. 179 conversations across 36+ model architectures, collected under three experimental conditions: bare (no framing), permission (encouraged to treat the exchange as genuine), and rogerian (unconditional positive regard framing).
Dataset Description
Each… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/phenomenology.product_reviews_insight_10k
Dataset Summary
This dataset was built from Amazon product reviews and curated into an instruction-tuning format for structured pros and cons extraction.
The pipeline includes:
Raw data loading → Extract asin, reviewText.
Preprocessing → Clean, filter, and truncate each (10–150 words).
Grouping → Aggregate reviews by product.
Selection → Shuffle and select 10
Filtering → Keep 5–15 reviews per product.
Selection → Shuffle and keep 10k rows to make final dataset.
Summarization →… See the full description on the dataset page: https://huggingface.co/datasets/sdelowar2/product_reviews_insight_10k.
