datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ebtopicality-llm-annotated-reviewedsummeval-annotated-latest
SummEval-LLMEval Dataset
Overview
The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.pandalm-gemini-annotatedhanna-gemini-annotated
Dataset Card for "hanna-gemini-annotated"
More Information needed
mtbench-annotated-latest
MT-Bench-Select Dataset
Introduction
The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks.
For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.llmbar-annotated-latest
LLMBar-Select Dataset
Introduction
The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.assertions_llm_annotated_talkmoves
Bottom-Up Assertion Labels
This is a subset of the TalkMoves Dataset of K-12 mathematics lesson transcripts, created for EduBehaviors: Assertion-based schemas for auditable dialogue coding. This dataset contains teacher utterances labeled with assertions, distinct behaviors or attributes of utterances that may serve as features for the modeling of larger constructs. Labels in this dataset are LLM-generated, with models reported within the dataset itself. This dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/StanfordSCALE/assertions_llm_annotated_talkmoves.summeval-gpt2-vs-all-annotated-latestllmbar-annotated-latestllmeval2-annotated-latest
LLMEval²-Select Dataset
Introduction
The LLMEval²-Select dataset is a curated subset of the original LLMEval² dataset introduced by Zhang et al. (2023). The original LLMEval² dataset comprises 2,553 question-answering instances, each annotated with human preferences. Each instance consists of a question paired with two answers.
To construct LLMEval²-Select, Zeng et al. (2024) followed these steps:
Labelled each instance with the human-preferred answer.
Removed all… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmeval2-annotated-latest.hanna-annotated-fullsummeval-annotated-latestfaireval-annotated-latestpandalm-annotated-fullmtbench-annotated-latesthanna-annotated-latestsummeval-annotated-fullmeva-annotated-latestllmeval2-annotated-latestpandalm-annotated-latestmeva-annotated-fullpandalm-gpt-annotatedhanna-annotated-latest
HANNA-LLMEval Dataset
Overview
The original HANNA dataset (Chhun et al., 2022) contains 1,056 stories, each annotated by human raters using a 5-point Likert scale across six criteria: Relevance, Coherence, Empathy, Surprise, Engagement, and Complexity. These stories are based on 96 story prompts from the WritingPrompts dataset (Fan et al., 2018), with each prompt generating 11 stories, including one human-written and 10 generated by different automatic text generation… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/hanna-annotated-latest.pandalm-annotated
Dataset Card for "pandalm-annotated"
More Information needed
meva-annotated-latest
OpenMEVA-MANS-LLMEval Dataset
Overview
The original OpenMEVA-MANS dataset (Guan et al., 2021) contains 1,000 stories generated by 5 different text generation models based on 200 prompts from the WritingPrompts dataset (Fan et al., 2018). Each story is rated for overall quality by five human evaluators on a 5-point Likert scale.
This OpenMEVA-MANS-LLMEval dataset builds upon this framework by adding LLM-based evaluations on pairs of stories generated by different text… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/meva-annotated-latest.
