CoolFace
20 results

prometheus

prometheus-eval /k-browsecomp K-BrowseComp K-BrowseComp is a Korean version of BrowseComp: a web-browsing agent benchmark. Items are grounded in Korean contexts and require retrieving information across multiple Korean websites. The 300-question verified subset is entirely handcrafted by native Korean speakers and every item underwent thorough manual revision and validation. 📄 Paper: https://arxiv.org/abs/2606.02404 💻 Code: https://github.com/prometheus-eval/K-BrowseComp Subsets… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/k-browsecomp.textquestion-answeringn<1K9 likes1.3k downloads4mo agoHugging Faceprometheus-eval /Feedback-Collection Dataset Card Dataset Summary The Feedback Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.\ Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\ In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Feedback-Collection.texttext-generation10K<n<100K120 likes557 downloads3y agoHugging Facechargoddard /WebInstructSub-prometheus Dataset Card for WebInstructSub-prometheus This dataset has been created with distilabel. Dataset Summary TIGER-Lab/WebInstructSub evaluated for logical and effective reasoning using prometheus-7b-v2.0. This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/WebInstructSub-prometheus.text1M<n<10M25 likes527 downloads2y agoHugging Faceprometheus-eval /peerreview-bench PeerReview Bench CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/ Repository:https://github.com/prometheus-eval/cmu-paper-reviewer Paper:https://arxiv.org/abs/2605.20668 Point of Contact:seungone@kaist.ac.kr Expert-annotated review items from scientific papers, organized for three complementary evaluation tasks. All data in this dataset is intended for evaluation, not training. All configs reference a shared, deduplicated file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.tabulartext-classification10K<n<100K3 likes395 downloads4mo agoHugging Faceprometheus-eval /BiGGen-Bench BIGGEN-Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models Dataset Description BIGGEN-Bench (BiG Generation Benchmark) is a comprehensive evaluation benchmark designed to assess the capabilities of large language models (LLMs) across a wide range of tasks. This benchmark focuses on free-form text generation and employs fine-grained, instance-specific evaluation criteria. Key Features: Purpose: To evaluate LLMs on diverse capabilities… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/BiGGen-Bench.texttext-generationn<1K17 likes363 downloads1y agoHugging Faceprometheus-eval /BiGGen-Bench-Results BIGGEN-Bench Evaluation Results Dataset Description This dataset contains the evaluation results for various language models on the BIGGEN-Bench (BiG Generation Benchmark). It provides comprehensive performance assessments across multiple capabilities and tasks. Key Features Evaluation results for 103 language models Scores across 9 different capabilities Results from multiple evaluator models (GPT-4, Claude-3-Opus, Prometheus-2) Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/BiGGen-Bench-Results.tabular10K<n<100K12 likes359 downloads2y agoHugging Face