datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
resultsnonmyopia_resultsresultsevaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}llm_pt_leaderboard_raw_resultsreward-bench-results
Results for Holisitic Evaluation of Reward Models (HERM) Benchmark
Here, you'll find the raw scores for the HERM project.
The repository is structured as follows.
├── best-of-n/ <- Nested directory for different completions on Best of N challenge
| ├── alpaca_eval/ └── results for each reward model
| | ├── tulu-13b/{org}/{model}.json
| | └── zephyr-7b/{org}/{model}.json
| └── mt_bench/
|… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench-results.resultsresultsorganic_query_results_datasetLiveHouse-TS-resultsleaderboard-resultsresults
HAKARI-Bench Results
This dataset stores raw benchmark result artifacts generated by HAKARI-Bench.
Raw results: per-task JSON (.xz) result files measured by HAKARI-bench.
Leaderboard: https://huggingface.co/spaces/hakari-bench/leaderboard
GitHub repository: https://github.com/hakari-bench/hakari-bench
Contributing official model results: follow the new model evaluation workflow to evaluate a model and submit results for HAKARI-Bench review:… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/results.p2-etf-levy-stable-resultseval-resultsresultsslo-rlvr-resultsneedle-resultsarena-resultsThis dataset contains the saved results from MTEB-Arena
resultsbtzsc-results
BTZSC Results
This repository stores model submissions for the BTZSC leaderboard.
BTZSC: A Benchmark for Zero-Shot Text Classification across Cross-Encoders, Embedding Models, Rerankers and LLMs.
Paper: https://openreview.net/forum?id=IxMryAz2p3
Eval harness: https://github.com/IliasAarab/btzsc
Leaderboard Space: https://huggingface.co/spaces/btzsc/btzsc-leaderboard
Benchmark summary:
22 English single-label datasets
4 task families: sentiment, topic, intent, emotion
Strict… See the full description on the dataset page: https://huggingface.co/datasets/btzsc/btzsc-results.vlm_resultslaal-resultssae_bench_resultsSWE-Bench-Verified-O1-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.lr_study_resultseval-resultsopen-asr-leaderboard-resultsda-code-evaluation-resultsp2-etf-hrp-allocator-resultsresults
Open LLM Leaderboard Results
This repository contains the outcomes of your submitted models that have been evaluated through the Open LLM Leaderboard. Our goal is to shed light on the cutting-edge Large Language Models (LLMs) and chatbots, enabling you to make well-informed decisions regarding your chosen application.
Evaluation Methodology
The evaluation process involves running your models against several benchmarks from the Eleuther AI Harness, a unified framework for… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/results.
