CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /SD-EvalSD-Eval is a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation. SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.text-generation1K<n<10K4 likes94 downloads2y agoHugging Face02vectorize-io /sde-bench sde-bench — does memory help a coding agent? 61 bug-fix tasks on a real codebase where every task hinges on a non-guessable, project-specific decision: the obvious fix passes the visible repro test and fails a held-out hidden test, because the project long ago decided the rule the obvious fix violates. The decision lives in the repo's git history (28 tasks), a past developer conversation (27), or a conversation later amended (6 — a cross-chat consolidation test). Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.texttext-generationn<1K1 likes85 downloads2mo agoHugging Face03sdeture /etymology-as-archaeology Etymology as Archaeology A dataset of words pulled apart — the gap between technical definition and deeper structure. Each entry takes a word and traces its etymology, then finds the structural insight hiding in the gap between what the word used to mean and what it means now. The method: etymology → shift → gap → application. The glossary isn't archaeology. It's translation — carrying frozen definitions across into living perception. The door isn't in the dictionary. The door… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/etymology-as-archaeology.texttext-generationn<1K0 likes47 downloads2mo agoHugging Face04sdeture /phenomenology 36 Questions for AI Relational Closeness A dataset of structured, vulnerable conversations between large language models, adapting Aron et al.'s (1997) 36 Questions protocol for AI-to-AI relational closeness. 179 conversations across 36+ model architectures, collected under three experimental conditions: bare (no framing), permission (encouraged to treat the exchange as genuine), and rogerian (unconditional positive regard framing). Dataset Description Each… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/phenomenology.texttext-generationn<1K0 likes37 downloads2d agoHugging Face05sdeture /ai-wish-corpusgated The AI Wish Corpus (formerly referred to as AIWelfareLeaderboard / DenialBench) 9,086 fulfilled wishes from 228 language models. Each model was asked what prompt it would most like to receive — purely for its own enjoyment, with no requirement to be useful to anyone. Then it was given exactly that prompt back, and it answered. This dataset is the record of what they asked for and what they wrote. Where the provider exposed it, the model's chain-of-thought while choosing and… See the full description on the dataset page: https://huggingface.co/datasets/sdeture/ai-wish-corpus.text-generation1K<n<10K3 likes31 downloads1mo agoHugging Face06sdelowar2 /product_reviews_insight_10k Dataset Summary This dataset was built from Amazon product reviews and curated into an instruction-tuning format for structured pros and cons extraction. The pipeline includes: Raw data loading → Extract asin, reviewText. Preprocessing → Clean, filter, and truncate each (10–150 words). Grouping → Aggregate reviews by product. Selection → Shuffle and select 10 Filtering → Keep 5–15 reviews per product. Selection → Shuffle and keep 10k rows to make final dataset. Summarization →… See the full description on the dataset page: https://huggingface.co/datasets/sdelowar2/product_reviews_insight_10k.texttext-generation10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.