open-weight
control-arena-persistent-state-glm-openweightdetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.automationbench600-open-weight-runs
AutomationBench 600 — raw run artifacts (open-weight models)
Complete raw exports from running the public 600-task
AutomationBench (Zapier, v1.0.5, --toolset api)
on open-weight models served locally with SGLang 0.5.12 on 8x H200.
Each .json.gz is the untouched --export-json output: run metadata, summary, and one record per
task including every message of the trajectory, the final environment state, and per-assertion
results.
File
Model
Pass rate
Partial credit… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/automationbench600-open-weight-runs.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted.evalap-comparing-openweight-with-patronusaiglider-114
Comparing openweight with PatronusAI/glider (ID: 114)
Comparing openweight Albert-API with specific judge PatronusAI/glider
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-patronusaiglider-114.
