datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drlc-leaderboard-dataopen-asr-leaderboard-resultsmmlu_pro_leaderboard_submissiondrlc-leaderboard-dataleaderboard_longformleaderboard_datahhem_leaderboard_datasetsbergson-wikitext-gpt2-leaderboard-bank
bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2)
Everything behind the numbers on the bergson leaderboard,
for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard.
path
what it is
bank/
the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv)
random/retrained/{base,subset_0..99}
the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.oruk-bench-leaderboard
oruk-bench leaderboard
Results for 64 speech emotion recognition systems measured on one held-out multilingual
evaluation with a single scoring implementation: open checkpoints, closed APIs, audio LLMs, and
text-only baselines, all on the same protocol.
This dataset is the results table, not the audio. The evaluation clips are assembled from
several emotional-speech corpora whose licences differ, so they are not redistributable; the
benchmark card documents
provenance and how… See the full description on the dataset page: https://huggingface.co/datasets/oruk/oruk-bench-leaderboard.ocr-leaderboard
OmniAI OCR Leaderboard
A comprehensive leaderboard comparing OCR and data extraction performance across traditional OCR providers and multimodal LLMs, such as gpt-4o and gemini-2.0. The dataset includes full results from testing 9 providers on 1,000 pages each.
Benchmark Results (Feb 2025) | Source Code
hivex-leaderboard-dataleaderboard-test
Benchmark Leaderboard
Αυτός είναι ο πίνακας αποτελεσμάτων.
tokenizer-leaderboard
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): en
License: mit
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/tokenizer-leaderboard.cwm-workout-leaderboard-datadrlc-leaderboard-dataleaderboard_resultsleaderboard-datasetscience_leaderboard_submissionThis dataset contains the results used for Science Leaderboard
open-llm-leaderboard-eda
🤖 Open LLM Leaderboard – Exploratory Data Analysis
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA)
of the Open LLM Leaderboard dataset from HuggingFace.
The goal is to understand what factors predict the overall benchmark
performance of open-source Large Language Models (LLMs).
The analysis is based on a dataset containing 4,575 LLM evaluation
records, including model size, training type, architecture,
and scores across 6 standardized… See the full description on the dataset page: https://huggingface.co/datasets/razsarusi/open-llm-leaderboard-eda.membench_leaderboard_submissionLLMs-Turkish-TEOG-Leaderboard
TEOG Scores Leaderboard
Welcome to the TEOG Scores Leaderboard! This repository contains the results of evaluating various large language models (LLMs) on the TEOG (Temel Eğitimden Ortaöğretime Geçiş) exam dataset. The TEOG exam is a standardized test in Turkey used for high school admissions, and this dataset provides a benchmark for assessing the performance of LLMs in Turkish educational tasks. Please remember that full score for TEOG is 500 points.
More Models Are… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/LLMs-Turkish-TEOG-Leaderboard.rl-leaderboard-requestsLLM-KG-Bench-LeaderboardLeaderboard for RDF Knowledge Graph(KG) related capabilities of Large Language Models(LLMs) as generated with the LLM-KG-Bench framework.
Results for more than 20 RDF related tasks are collected for more than 40 LLMs.
The leaderboard contains summarized results:
board_combined_scores_short.csv: the most concise summary, listing for each LLM combined scores in the RDF (R) and SPARQL (S) handling categories, estimating read(R) and write(W) capabilities for syntax(syn) and semantic(sem). If a… See the full description on the dataset page: https://huggingface.co/datasets/lpmeyer/LLM-KG-Bench-Leaderboard.ko_leaderboard한국어 리더보드 학습에서 사용된 데이터 중 약 8,000건에 대해서
공개합니다. 데이터 생성 시 도움이 되기를 바랍니다.
감사합니다.
shadropen-asr-leaderboard-evals-allarlbench-leaderboard-results
ARLBench Leaderboard Data
This is the data repository for the ARLBench leaderboard. If you want to add your own runs to the leaderboard, please make a PR here with the following content:
The data file(s) with your data. There's a template available you can fill in. If you add data for all algorithms, you can use a single file. If you're only adding for a subset, add separate files per algorithm.
The extended data split. This means extending the ReadMe.md config. If you added a file… See the full description on the dataset page: https://huggingface.co/datasets/autorl-org/arlbench-leaderboard-results.ai-adoption-leaderboards
AI Adoption Leaderboards (US Geography & Industry)
6,822 pre-computed leaderboards ranking AI adoption across all 50 US states, 300+ metros, 1,200+ cities, and 100+ NAICS industries — with company counts, average scores, and top companies per segment.
Rows: 6,822
Source: Derived from the AI Adoption Index (US Companies)
Methodology + interactive explorer: https://meoadvisors.com/ai-opportunities/leaderboard/
Part of: the open AI Workforce Data collection by Meo Advisors… See the full description on the dataset page: https://huggingface.co/datasets/Meo-Advisors/ai-adoption-leaderboards.
