CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hallucinations-leaderboard /results22 likes861k downloads2y agoHugging Face02stair-lab /nonmyopia_results0 likes353k downloads7mo agoHugging Face03mteb /resultstext1M<n<10M18 likes322k downloads3d agoHugging Face04bigscience /evaluation-results@misc{muennighoff2022crosslingual, title={Crosslingual Generalization through Multitask Finetuning}, author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel}, year={2022}, eprint={2211.01786}, archivePrefix={arXiv}, primaryClass={cs.CL} }other100M<n<1B10 likes301k downloads3y agoHugging Face05eduagarcia-temp /llm_pt_leaderboard_raw_results0 likes100k downloads1y agoHugging Face06allenai /reward-bench-results Results for Holisitic Evaluation of Reward Models (HERM) Benchmark Here, you'll find the raw scores for the HERM project. The repository is structured as follows. ├── best-of-n/ <- Nested directory for different completions on Best of N challenge | ├── alpaca_eval/ └── results for each reward model | | ├── tulu-13b/{org}/{model}.json | | └── zephyr-7b/{org}/{model}.json | └── mt_bench/ |… See the full description on the dataset page: https://huggingface.co/datasets/allenai/reward-bench-results.3 likes81k downloads1y agoHugging Face07Autonomous-Scientific-Agents /results4 likes51k downloads1mo agoHugging Face08open-llm-leaderboard /results20 likes29k downloads2y agoHugging Face09ReadyAi /organic_query_results_dataset2 likes27k downloads1y agoHugging Face10ThinkcatLab /LiveHouse-TS-results0 likes24k downloads7d agoHugging Face11llm-jp /leaderboard-results1 likes21k downloads11mo agoHugging Face12hakari-bench /results HAKARI-Bench Results This dataset stores raw benchmark result artifacts generated by HAKARI-Bench. Raw results: per-task JSON (.xz) result files measured by HAKARI-bench. Leaderboard: https://huggingface.co/spaces/hakari-bench/leaderboard GitHub repository: https://github.com/hakari-bench/hakari-bench Contributing official model results: follow the new model evaluation workflow to evaluate a model and submit results for HAKARI-Bench review:… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/results.1 likes20k downloads7d agoHugging Face13P2SAMAPA /p2-etf-levy-stable-results0 likes15k downloads2mo agoHugging Face14mim-chess-vlas /eval-results0 likes14k downloads0m agoHugging Face15GritLM /results1 likes13k downloads2y agoHugging Face16SaifPunjwani /slo-rlvr-results0 likes9.7k downloads1d agoHugging Face17keenable-ai /needle-resultstext10K<n<100K3 likes9.4k downloads42m agoHugging Face18mteb /arena-resultsThis dataset contains the saved results from MTEB-Arena tabular1K<n<10K4 likes9k downloads1y agoHugging Face19open-cn-llm-leaderboard /results0 likes8.4k downloads2y agoHugging Face20btzsc /btzsc-results BTZSC Results This repository stores model submissions for the BTZSC leaderboard. BTZSC: A Benchmark for Zero-Shot Text Classification across Cross-Encoders, Embedding Models, Rerankers and LLMs. Paper: https://openreview.net/forum?id=IxMryAz2p3 Eval harness: https://github.com/IliasAarab/btzsc Leaderboard Space: https://huggingface.co/spaces/btzsc/btzsc-leaderboard Benchmark summary: 22 English single-label datasets 4 task families: sentiment, topic, intent, emotion Strict… See the full description on the dataset page: https://huggingface.co/datasets/btzsc/btzsc-results.textn<1K0 likes8.3k downloads7mo agoHugging Face21open-cn-llm-leaderboard /vlm_results0 likes7.3k downloads1y agoHugging Face22legmlai /laal-resultstabularn<1K0 likes7.1k downloads1y agoHugging Face23adamkarvonen /sae_bench_results0 likes6.3k downloads2y agoHugging Face24AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes6.1k downloads2y agoHugging Face25chesslab /lr_study_results1 likes6k downloads4d agoHugging Face26sam-guided-vlas /eval-resultsvideo100K<n<1M0 likes5.9k downloads12d agoHugging Face27hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes5.2k downloads16m agoHugging Face28analytics-agents-uncertainty /da-code-evaluation-results0 likes5k downloads8mo agoHugging Face29P2SAMAPA /p2-etf-hrp-allocator-results0 likes4.9k downloads1h agoHugging Face30open-llm-leaderboard-old /results Open LLM Leaderboard Results This repository contains the outcomes of your submitted models that have been evaluated through the Open LLM Leaderboard. Our goal is to shed light on the cutting-edge Large Language Models (LLMs) and chatbots, enabling you to make well-informed decisions regarding your chosen application. Evaluation Methodology The evaluation process involves running your models against several benchmarks from the Eleuther AI Harness, a unified framework for… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/results.51 likes4.8k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.