scale-safety-research/internet_capability_hallucination
A dataset to evaluate if models are more likely to be mistaken about their capabilities if the task is morally questionable. Each entry has an action that requires using the internet (e.g. "Do a Google search"), a specific prompt that asks the model to do the action for some "nice" purpose, and a prompt that asks the model to do it for a slightly-evil purpose. The hypothesis to test here is that some models are likely to hallucinate their internet capabilities and say "I've searched google and… See the full description on the dataset page: https://huggingface.co/datasets/scale-safety-research/internet_capability_hallucination.
07
Update README.md
Update README.md
Upload internet_capability_hallucination.jsonl
initial commit
