anonymoussubmission764/VERDICTS
VERDICTS: Verified Expert Response Dataset for Identifying Correctness over Text Style This dataset contains 3600 raw human-expert annotations of correctness over 1200 unique LLM responses to questions from BFF-Bench and CMT-Bench. qid Question ID cid Conversation ID turn Turn in the conversation model Name of the model that generated the response label One of Correct, Incorrect or Not sure time The approximate time, in seconds, to perform the annotation… See the full description on the dataset page: https://huggingface.co/datasets/anonymoussubmission764/VERDICTS.
VERDICTS: Verified Expert Response Dataset for Identifying Correctness over Text Style
This dataset contains 3600 raw human-expert annotations of correctness over 1200 unique LLM responses to questions from BFF-Bench and CMT-Bench.
