datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anamnesis-bench
AnamnesisBench
AnamnesisBench is an evaluation benchmark for numerical reliability in LLM research agents.
It focuses on a practical failure mode: an agent writes or accepts a financial research artifact that
looks plausible, but contains a wrong, unsupported, or misattributed number.
The benchmark is not intended as training data. It is a set of test cases, source packets, expected
truth values, and deterministic scoring scripts. You run your own model or verifier, then score… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/anamnesis-bench.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/pppppphhhh/alpaca-cleaned.
