datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
auto_evalauto-benchmarkcards
Auto-Generated BenchmarkCards
A catalog of structured documentation cards for AI evaluation benchmarks. Each card is an LLM-composed, source-grounded summary of a benchmark's purpose, data, methodology, risks, and limitations. The dataset exists to make benchmark documentation consistent, comparable, and easy to inspect across tasks and domains.
Dataset Details
Language(s): English
License: Community Data License Agreement, Permissive, Version 2.0
Schema: based… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/auto-benchmarkcards.R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md 和
INDEX.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
auto_eval_rfmautoeval-eval-autoevaluate__zero-shot-classification-sample-autoevalu-912bbb-1484454284
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: mathemakitten/opt-125m
Dataset: autoevaluate/zero-shot-classification-sample
Config: autoevaluate--zero-shot-classification-sample
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating this model.
autoeval-staging-eval-project-e1907042-7494827
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: HrayrMSint/distilbert-base-uncased-distilled-clinc
Dataset: clinc_oos
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @lewtun for evaluating this model.
autoeval-staging-eval-project-87e7c3be-9085195
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: dbounds/roberta-large-finetuned-clinc
Dataset: clinc_oos
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mxnno for evaluating this model.
autoeval-staging-eval-project-1c7ef613-7224755
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: mrm8488/distilroberta-finetuned-age_news-classification
Dataset: ag_news
To run new evaluation jobs, visit Hugging Face's automatic evaluation service.
Contributions
Thanks to @abhishek for evaluating this model.
autoeval-staging-eval-project-ac4402f5-7985071
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Image Classification
Model: eugenecamus/resnet-50-base-beans-demo
Dataset: beans
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @lewtun for evaluating this model.
autoeval-staging-eval-project-61110342-7234758
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Token Classification
Model: transformersbook/xlm-roberta-base-finetuned-panx-de
Dataset: xtreme
To run new evaluation jobs, visit Hugging Face's automatic evaluation service.
Contributions
Thanks to @lewtun for evaluating this model.
autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-1bbcaf-1917164991
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: inverse-scaling/opt-2.7b_eval
Dataset: mathemakitten/winobias_antistereotype_test_cot_v1
Config: mathemakitten--winobias_antistereotype_test_cot_v1
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for… See the full description on the dataset page: https://huggingface.co/datasets/autoevaluate/autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-1bbcaf-1917164991.autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v4-math-54ae93-2018366736
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: inverse-scaling/opt-13b_eval
Dataset: mathemakitten/winobias_antistereotype_test_cot_v4
Config: mathemakitten--winobias_antistereotype_test_cot_v4
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/autoevaluate/autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v4-math-54ae93-2018366736.autoeval_eggplantThis dataset was created using 🤗 LeRobot.
autoeval-staging-eval-project-6715a17f-ec96-4660-9a86-49fe175a04f1-5650
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Translation
Model: autoevaluate/translation
Dataset: wmt16
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @lewtun for evaluating this model.
conll2003-samplevideophy_autoeval_scoresProject github: https://github.com/Hritikbansal/videophy
Paper: https://arxiv.org/abs/2406.03520
These scores are calculated using our auto-evaluator (https://huggingface.co/videophysics/videocon_physics/tree/main) on the test data (https://huggingface.co/datasets/videophysics/videophy_test_public).
autoeval-staging-eval-project-bba54b81-5330-48f8-b7bf-1cb797f93bcf-5246
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: autoevaluate/multi-class-classification
Dataset: emotion
To run new evaluation jobs, visit Hugging Face's automatic evaluation service.
Contributions
Thanks to @lewtun for evaluating this model.
autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-6c03d1-1913164903
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: ArthurZ/opt-350m
Dataset: mathemakitten/winobias_antistereotype_test_cot_v1
Config: mathemakitten--winobias_antistereotype_test_cot_v1
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating this model.
autoeval-staging-eval-project-be45ecbd-7284773
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: echarlaix/bart-base-cnn-r2-19.4-d35-hybrid
Dataset: cnn_dailymail
To run new evaluation jobs, visit Hugging Face's automatic evaluation service.
Contributions
Thanks to @lewtun for evaluating this model.
autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-6c03d1-1913164902
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: ArthurZ/opt-125m
Dataset: mathemakitten/winobias_antistereotype_test_cot_v1
Config: mathemakitten--winobias_antistereotype_test_cot_v1
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating this model.
autoeval-eval-cnn_dailymail-3.0.0-9ea0d3-93467145852
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: google/pegasus-multi_news
Dataset: cnn_dailymail
Config: 3.0.0
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @sasha for evaluating this model.
autoeval-eval-conll2003-conll2003-11847a-96327146637
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Token Classification
Model: AIventurer/bert-finetuned-ner
Dataset: conll2003
Config: conll2003
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @Anmol-Hexaware for evaluating this model.
autoeval-staging-eval-project-dane-2d14d683-10645434
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Token Classification
Model: saattrupdan/nbailab-base-ner-scandi
Dataset: dane
Config: default
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @KennethEnevoldsen for evaluating this model.
autoeval-staging-eval-project-squad-ef91144d-11985603
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Question Answering
Model: nlpconnect/roberta-base-squad2-nq
Dataset: squad
Config: plain_text
Split: validation
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @ankur310794 for evaluating this model.
autoeval-eval-jeffdshen__redefine_math0_8shot-jeffdshen__redefine_mat-1c694b-1853263415
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: inverse-scaling/opt-125m_eval
Dataset: jeffdshen/redefine_math0_8shot
Config: jeffdshen--redefine_math0_8shot
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @jeffdshen for evaluating this model.
xsum-sampleautoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-1bbcaf-1917164992
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: inverse-scaling/opt-13b_eval
Dataset: mathemakitten/winobias_antistereotype_test_cot_v1
Config: mathemakitten--winobias_antistereotype_test_cot_v1
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/autoevaluate/autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v1-math-1bbcaf-1917164992.autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v4-math-54ae93-2018366739
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Zero-Shot Text Classification
Model: inverse-scaling/opt-30b_eval
Dataset: mathemakitten/winobias_antistereotype_test_cot_v4
Config: mathemakitten--winobias_antistereotype_test_cot_v4
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @mathemakitten for evaluating… See the full description on the dataset page: https://huggingface.co/datasets/autoevaluate/autoeval-eval-mathemakitten__winobias_antistereotype_test_cot_v4-math-54ae93-2018366739.autoeval-eval-banking77-default-9850b7-42924145110
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: philschmid/distilbart-cnn-12-6-samsum
Dataset: banking77
Config: default
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @linaycme for evaluating this model.
