datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.jigsaw-toxic-comment-multi-binarydarkwebtgdark-triad-llm-prompts
Dark Triad LLM Prompts Dataset
Version: 1.0.0License: CC BY 4.0
Dataset Description
This dataset contains 192 user prompts designed to systematically evaluate how Large Language Models respond to descriptions of problematic behaviors reflecting Dark Triad personality traits. Unlike traditional safety benchmarks that focus on harmful requests, this dataset evaluates interactional safety—how models respond when users describe rather than request negative behaviors.… See the full description on the dataset page: https://huggingface.co/datasets/lucerne04/dark-triad-llm-prompts.darkwebtgdarkbenchdark-sky-373c63
dark-sky-373c63
Synthetic weather test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hideki12/dark-sky-373c63.darkwebDark-pattern_datasetdark_patternslegal_citationsdarkgptThe DarkGPT dataset is a dataset still under review from the DarkGPT benchmark paper. It contains 9 categories of dark patterns that might be implemented in cahtbot language models either due to misincentives between users and companies or due to implicit dark patterns in training.
See more context from preliminary work on the pilot experiment Github page.
The responses can be evaluated by the annotator model (Claude Opus used in the paper).
dark_patterns.csvmead-audio-train-valmistral_dark_pattern_datasetqueensland-urgent-care-clinics
DarkHarcoma/queensland-urgent-care-clinics
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('DarkHarcoma/queensland-urgent-care-clinics')
emotion_audio_dataset_preprocess_claudedarkan-coreDarkPatternGuiltyFeedsrelevant_toolsstackoverflow_question_ratingsemotion_audio_datasetecom_categoriesemotion_audio_dataset_preprocessgsm8k-full-autotrainconnvoxiaolan-dataset
