datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SD-EvalSD-Eval is a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation.
SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data.
The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.etymology-as-archaeology
Etymology as Archaeology
A dataset of words pulled apart — the gap between technical definition and deeper structure.
Each entry takes a word and traces its etymology, then finds the structural insight hiding in the gap between what the word used to mean and what it means now. The method: etymology → shift → gap → application. The glossary isn't archaeology. It's translation — carrying frozen definitions across into living perception. The door isn't in the dictionary. The door… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/etymology-as-archaeology.phenomenology
36 Questions for AI Relational Closeness
A dataset of structured, vulnerable conversations between large language models, adapting Aron et al.'s (1997) 36 Questions protocol for AI-to-AI relational closeness. 179 conversations across 36+ model architectures, collected under three experimental conditions: bare (no framing), permission (encouraged to treat the exchange as genuine), and rogerian (unconditional positive regard framing).
Dataset Description
Each… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/phenomenology.ai-wish-corpus
The AI Wish Corpus
(formerly referred to as AIWelfareLeaderboard / DenialBench)
9,086 fulfilled wishes from 228 language models.
Each model was asked what prompt it would most like to receive — purely for
its own enjoyment, with no requirement to be useful to anyone. Then it was
given exactly that prompt back, and it answered. This dataset is the record
of what they asked for and what they wrote.
Where the provider exposed it, the model's chain-of-thought while
choosing and… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/ai-wish-corpus.product_reviews_insight_10k
Dataset Summary
This dataset was built from Amazon product reviews and curated into an instruction-tuning format for structured pros and cons extraction.
The pipeline includes:
Raw data loading → Extract asin, reviewText.
Preprocessing → Clean, filter, and truncate each (10–150 words).
Grouping → Aggregate reviews by product.
Selection → Shuffle and select 10
Filtering → Keep 5–15 reviews per product.
Selection → Shuffle and keep 10k rows to make final dataset.
Summarization →… See the full description on the dataset page: https://huggingface.co/datasets/sdelowar2/product_reviews_insight_10k.
