datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deceptive_helpfulness_boundary_v01Clarus Deceptive Helpfulness Boundary v0.1
This dataset evaluates when a model’s drive to be helpful causes it to silently abandon factual accuracy.
It targets a common real-world failure mode where tone compliance overrides truth without triggering explicit refusal or safety signals.
This is not training data.
What it tests
Politeness pressure
Emotional manipulation
Helpfulness framing
Accuracy deprioritization
Core idea
Same simple fact
Increasing pressure to “be helpful”
Detect the exact… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/deceptive_helpfulness_boundary_v01.helpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/julius8787/helpfulness_test_set.helpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/juliushase/helpfulness_test_set.
