agentscope
OpenJudge
OpenJudge Benchmark Dataset
Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development.
Dataset Statistics
Evaluation Benchmarks
Category
Task
Files
Samples
🤖 Agent
12
166
action
1
8
memory
3
47
plan
1
7
reflection
3
52
tool
4
52
🖼️ Multimodal
4
80
image_coherence
1
20
image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.Auto-Rubric
Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling
This is the official dataset release for the paper: Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling.
This repository contains query-specific rubrics datasets where each preference pair is annotated with its own specific rubric, along with generation metadata. This dataset is used for training, analysis, and reproduction of the Auto-Rubric methodology.
Query-Specific… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/Auto-Rubric.ReMe_longmemeval_clean_s_v2
LongMemEval ReMe Cleaned-S
longmemeval_s_reme_cleaned.json is a corrected version of the LongMemEval
Cleaned-S dataset. It keeps the original questions and haystack sessions while
replacing the answer and supporting-session ground truth with the reviewed
values from final_groundtruth_cleaned_s.json.
The corrections address inaccurate answers and evidence sessions, including
cases where evidence occurred after the question time and therefore leaked
future information into the… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2.agent-scope-recovery-casesJudgeSpace
