Auto-eval
auto_evalauto-benchmarkcards
Auto-Generated BenchmarkCards
A catalog of structured documentation cards for AI evaluation benchmarks. Each card is an LLM-composed, source-grounded summary of a benchmark's purpose, data, methodology, risks, and limitations. The dataset exists to make benchmark documentation consistent, comparable, and easy to inspect across tasks and domains.
Dataset Details
Language(s): English
License: Community Data License Agreement, Permissive, Version 2.0
Schema: based… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/auto-benchmarkcards.R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md 和
INDEX.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
auto_eval_rfmautoeval-eval-autoevaluate__zero-shot-classification-sample-autoevalu-912bbb-1484454284
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: mathemakitten/opt-125m
Dataset: autoevaluate/zero-shot-classification-sample
Config: autoevaluate--zero-shot-classification-sample
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating this model.
