CoolFace
12 results

llm-annotated

JPPOL-AI /ebtopicality-llm-annotated-reviewedtextn<1K0 likes71 downloads22d agoHugging Facebay-calibration-llm-evaluators /summeval-annotated-latest SummEval-LLMEval Dataset Overview The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.tabular1K<n<10K0 likes51 downloads2y agoHugging Facellm-aes /pandalm-gemini-annotatedtabular1K<n<10K0 likes49 downloads3y agoHugging Facellm-aes /hanna-gemini-annotated Dataset Card for "hanna-gemini-annotated" More Information needed tabular10K<n<100K0 likes47 downloads3y agoHugging Facebay-calibration-llm-evaluators /mtbench-annotated-latest MT-Bench-Select Dataset Introduction The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks. For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.tabular1K<n<10K0 likes34 downloads2y agoHugging Facebay-calibration-llm-evaluators /llmbar-annotated-latest LLMBar-Select Dataset Introduction The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.tabular1K<n<10K0 likes33 downloads2y agoHugging Face