CoolFace
Datasetpublic

anonymoussubmission764/VERDICTS

VERDICTS: Verified Expert Response Dataset for Identifying Correctness over Text Style This dataset contains 3600 raw human-expert annotations of correctness over 1200 unique LLM responses to questions from BFF-Bench and CMT-Bench. qid Question ID cid Conversation ID turn Turn in the conversation model Name of the model that generated the response label One of Correct, Incorrect or Not sure time The approximate time, in seconds, to perform the annotation… See the full description on the dataset page: https://huggingface.co/datasets/anonymoussubmission764/VERDICTS.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes2downloads
Dataset Card

VERDICTS: Verified Expert Response Dataset for Identifying Correctness over Text Style

This dataset contains 3600 raw human-expert annotations of correctness over 1200 unique LLM responses to questions from BFF-Bench and CMT-Bench.

qidQuestion ID
cidConversation ID
turnTurn in the conversation
modelName of the model that generated the response
labelOne of Correct, Incorrect or Not sure
timeThe approximate time, in seconds, to perform the annotation
datasetOne of bffbench or mtbmr (MT-Bench Math and Reasoning)
convjson representation of the conversation