prometheus
k-browsecomp
K-BrowseComp
K-BrowseComp is a Korean version of BrowseComp: a web-browsing agent benchmark. Items are grounded in Korean contexts and require retrieving information across multiple Korean websites.
The 300-question verified subset is entirely handcrafted by native Korean speakers and every item underwent thorough manual revision and validation.
📄 Paper: https://arxiv.org/abs/2606.02404
💻 Code: https://github.com/prometheus-eval/K-BrowseComp
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/k-browsecomp.Feedback-Collection
Dataset Card
Dataset Summary
The Feedback Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.\
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Feedback-Collection.WebInstructSub-prometheus
Dataset Card for WebInstructSub-prometheus
This dataset has been created with distilabel.
Dataset Summary
TIGER-Lab/WebInstructSub evaluated for logical and effective reasoning using prometheus-7b-v2.0.
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/WebInstructSub-prometheus.peerreview-bench
PeerReview Bench
CMU Paper Reviewer:https://prometheus-eval.github.io/cmu-paper-reviewer/
Repository:https://github.com/prometheus-eval/cmu-paper-reviewer
Paper:https://arxiv.org/abs/2605.20668
Point of Contact:seungone@kaist.ac.kr
Expert-annotated review items from scientific papers, organized for three
complementary evaluation tasks. All data in this dataset is intended
for evaluation, not training. All configs reference a shared, deduplicated
file store (submitted_papers)… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/peerreview-bench.BiGGen-Bench
BIGGEN-Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models
Dataset Description
BIGGEN-Bench (BiG Generation Benchmark) is a comprehensive evaluation benchmark designed to assess the capabilities of large language models (LLMs) across a wide range of tasks. This benchmark focuses on free-form text generation and employs fine-grained, instance-specific evaluation criteria.
Key Features:
Purpose: To evaluate LLMs on diverse capabilities… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/BiGGen-Bench.BiGGen-Bench-Results
BIGGEN-Bench Evaluation Results
Dataset Description
This dataset contains the evaluation results for various language models on the BIGGEN-Bench (BiG Generation Benchmark). It provides comprehensive performance assessments across multiple capabilities and tasks.
Key Features
Evaluation results for 103 language models
Scores across 9 different capabilities
Results from multiple evaluator models (GPT-4, Claude-3-Opus, Prometheus-2)
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/BiGGen-Bench-Results.
