datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vicgalle__CarbonBeagle-11B-truthy-details
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.Qwen__Qwen2.5-7B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-7B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-Instruct-details.HuggingFaceH4__zephyr-7b-beta-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.Qwen__Qwen2.5-72B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.mistralai__Mistral-7B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1
The dataset is composed of 82 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 34 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.1-details.meta-llama__Meta-Llama-3-70B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.details_bestdive__SmolLM3-3B-SFT-Free-Course
Smol course SFT evaluation - Kay Zheng
Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation.
Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422).
Free Google Colab T4, no paid HF Jobs; cost 0.
lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12.
Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.deepseek-ai__deepseek-llm-7b-chat-details
Dataset Card for Evaluation run of deepseek-ai/deepseek-llm-7b-chat
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-llm-7b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-llm-7b-chat-details.01-ai__Yi-1.5-34B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-34B-Chat-details.mistralai__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.ai21labs__Jamba-v0.1-details
Dataset Card for Evaluation run of ai21labs/Jamba-v0.1
Dataset automatically created during the evaluation run of model ai21labs/Jamba-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ai21labs__Jamba-v0.1-details.01-ai__Yi-34B-details
Dataset Card for Evaluation run of 01-ai/Yi-34B
Dataset automatically created during the evaluation run of model 01-ai/Yi-34B
The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-34B-details.Deci__DeciLM-7B-details
Dataset Card for Evaluation run of Deci/DeciLM-7B
Dataset automatically created during the evaluation run of model Deci/DeciLM-7B
The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Deci__DeciLM-7B-details.bigcode__starcoder2-3b-details
Dataset Card for Evaluation run of bigcode/starcoder2-3b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.abacusai__Llama-3-Smaug-8B-details
Dataset Card for Evaluation run of abacusai/Llama-3-Smaug-8B
Dataset automatically created during the evaluation run of model abacusai/Llama-3-Smaug-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Llama-3-Smaug-8B-details.HelpingAI__Dhanishtha-Large-details
Dataset Card for Evaluation run of HelpingAI/Dhanishtha-Large
Dataset automatically created during the evaluation run of model HelpingAI/Dhanishtha-Large
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HelpingAI__Dhanishtha-Large-details.01-ai__Yi-1.5-9B-Chat-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B-Chat
The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-Chat-details.HuggingFaceH4__zephyr-7b-alpha-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-alpha
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-alpha
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-alpha-details.deepseek-ai__deepseek-moe-16b-chat-details
Dataset Card for Evaluation run of deepseek-ai/deepseek-moe-16b-chat
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-moe-16b-chat
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-moe-16b-chat-details.mistralai__Mistral-Large-Instruct-2411-details
Dataset Card for Evaluation run of mistralai/Mistral-Large-Instruct-2411
Dataset automatically created during the evaluation run of model mistralai/Mistral-Large-Instruct-2411
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Large-Instruct-2411-details.OpenAssistant__oasst-sft-1-pythia-12b-details
Dataset Card for Evaluation run of OpenAssistant/oasst-sft-1-pythia-12b
Dataset automatically created during the evaluation run of model OpenAssistant/oasst-sft-1-pythia-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenAssistant__oasst-sft-1-pythia-12b-details.JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details
Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3
Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.Qwen__Qwen2.5-32B-Instruct-details
Dataset Card for Evaluation run of Qwen/Qwen2.5-32B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-32B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-32B-Instruct-details.tiiuae__Falcon3-7B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-7B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-7B-Instruct-details.ibm-granite__granite-3.0-2b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-base-details.Aurel9__testmerge-7b-details
Dataset Card for Evaluation run of Aurel9/testmerge-7b
Dataset automatically created during the evaluation run of model Aurel9/testmerge-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Aurel9__testmerge-7b-details.EleutherAI__gpt-j-6b-details
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b
The dataset is composed of 111 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-j-6b-details.01-ai__Yi-1.5-9B-details
Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-details.deepseek-ai__DeepSeek-R1-Distill-Qwen-1.5B-details
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__DeepSeek-R1-Distill-Qwen-1.5B-details.
