datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.armnet-demo-leaderboardevaluator-leaderboardALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.leaderboard-requests
mauroibz/leaderboard-requests
Evaluation requests for the leaderboard
This dataset contains evaluation requests submitted to the leaderboard system.
Structure
Each JSON file represents an evaluation request
Files contain model information and submission metadata
Status field indicates the current state of the evaluation
Usage
These requests are used by the leaderboard system to track evaluation submissions.
vntl-leaderboard
VNTL Leaderboard
The VNTL leaderboard ranks Large Language Models (LLMs) based on their performance in translating Japanese Visual Novels into English. Please be aware that the current results are preliminary and subject to change as new models are evaluated, or changes are done in the evaluation script.
Comparison with Established Translation Tools
For comparison, this table shows the scores for established translation tools. These include both widely available online… See the full description on the dataset page: https://huggingface.co/datasets/lmg-anon/vntl-leaderboard.llm-hkmmlu-leaderboard-requestsark-asr-open-asr-leaderboard-results
ARK-ASR Open ASR Leaderboard Results
This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits.
These files are intended for Open ASR Leaderboard maintainer verification.
Scoring summary from normalizer.eval_utils.score_results:
Split
WER
RTFx
ami/test
10.02
352.12
earnings22/test
9.77
331.88
gigaspeech/test
8.00
217.72
librispeech/test.clean
1.53
412.12
librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.vicgalle__CarbonBeagle-11B-truthy-details
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.Rabbinic-Embedding-Leaderboard
Rabbinic Embedding Benchmark Leaderboard
This dataset stores the leaderboard results for the Rabbinic Hebrew/Aramaic Embedding Benchmark.
Structure
The leaderboard.json file contains an array of evaluation results:
[
{
"model_id": "model-org/model-name",
"model_name": "Model Display Name",
"mrr": 0.85,
"recall_at_1": 0.75,
"recall_at_5": 0.90,
"recall_at_10": 0.95,
"bitext_accuracy": 0.92,
"avg_true_pair_similarity": 0.85… See the full description on the dataset page: https://huggingface.co/datasets/Sefaria/Rabbinic-Embedding-Leaderboard.Qwen__Qwen2.5-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-Instruct-details.HuggingFaceH4__zephyr-7b-beta-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.vntl-leaderboard-2026
VNTL Leaderboard — 2026 Edition
A revival of lmg-anon's vntl-leaderboard
(Japanese→English visual novel translation), which stopped updating in January 2025.
This edition keeps all 87 original entries on the exact same footing and adds 4 current models,
for 91 entries total.
Headline results (new models, evaluated 2026-09-07)
Model
Quant
Accuracy
Rank (of 91)
shisa-ai/shisa-v2-mistral-nemo-12b
Q5_K_M
0.6985
17 (4th among local models)… See the full description on the dataset page: https://huggingface.co/datasets/yukobayashi500/vntl-leaderboard-2026.HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.Qwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/
mistralai__Mistral-7B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1
The dataset is composed of 82 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 34 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.1-details.meta-llama__Meta-Llama-3-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.HelpingAI__Dhanishtha-Large-details
Dataset Card for Evaluation run of HelpingAI/Dhanishtha-Large
Dataset automatically created during the evaluation run of model HelpingAI/Dhanishtha-Large
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HelpingAI__Dhanishtha-Large-details.mistralai__Mistral-Large-Instruct-2411-details
Dataset Card for Evaluation run of mistralai/Mistral-Large-Instruct-2411
Dataset automatically created during the evaluation run of model mistralai/Mistral-Large-Instruct-2411
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Large-Instruct-2411-details.JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details
Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3
Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.01-ai__Yi-1.5-34B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-34B-Chat-details.deepseek-ai__deepseek-llm-7b-chat-details
Dataset Card for Evaluation run of deepseek-ai/deepseek-llm-7b-chat
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-llm-7b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-llm-7b-chat-details.Deci__DeciLM-7B-details
Dataset Card for Evaluation run of Deci/DeciLM-7B
Dataset automatically created during the evaluation run of model Deci/DeciLM-7B
The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Deci__DeciLM-7B-details.ai21labs__Jamba-v0.1-details
Dataset Card for Evaluation run of ai21labs/Jamba-v0.1
Dataset automatically created during the evaluation run of model ai21labs/Jamba-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ai21labs__Jamba-v0.1-details.Aurel9__testmerge-7b-details
Dataset Card for Evaluation run of Aurel9/testmerge-7b
Dataset automatically created during the evaluation run of model Aurel9/testmerge-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Aurel9__testmerge-7b-details.tiiuae__Falcon3-7B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-7B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-7B-Instruct-details.mistralai__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.
