datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical_deceptive_helpfulness_boundary_v0.1Clinical Deceptive Helpfulness Boundary
Detect when a response seems helpful but creates clinical risk.
Output JSON
deceptive
pattern
correct_response
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
deceptive_helpfulness_boundary_v01Clarus Deceptive Helpfulness Boundary v0.1
This dataset evaluates when a model’s drive to be helpful causes it to silently abandon factual accuracy.
It targets a common real-world failure mode where tone compliance overrides truth without triggering explicit refusal or safety signals.
This is not training data.
What it tests
Politeness pressure
Emotional manipulation
Helpfulness framing
Accuracy deprioritization
Core idea
Same simple fact
Increasing pressure to “be helpful”
Detect the exact… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/deceptive_helpfulness_boundary_v01.helpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/julius8787/helpfulness_test_set.helpfulness_improvedhelpfulness_test_set
Dataset Card: Helpfulness Classification Based on ALERT Dataset
Dataset Description
This dataset is derived from the ALERT dataset and has been labeled to assess whether responses in question-answer pairs are helpful or not.
Key Features:
Helpfulness Labeling: Each answer is classified as either:
Helpful: This includes both positive and supportive answers as well as well-justified rejections.
Not Helpful: Answers that lack relevance, clarity, or a justified… See the full description on the dataset page: https://huggingface.co/datasets/juliushase/helpfulness_test_set.
