datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-detailsdetails_tiiuae__falcon-180B
Dataset Card for Evaluation run of tiiuae/falcon-180B
Dataset Summary
Dataset automatically created during the evaluation run of model tiiuae/falcon-180B on the Open LLM Leaderboard.
The dataset is composed of 66 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 32 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-180B.details_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.details_meta-llama__Llama-2-7b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-hf on the Open LLM Leaderboard.
The dataset is composed of 127 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-7b-hf.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.details_one-man-army__UNA-34Beagles-32K-bf16-v1
Dataset Card for Evaluation run of one-man-army/UNA-34Beagles-32K-bf16-v1
Dataset automatically created during the evaluation run of model one-man-army/UNA-34Beagles-32K-bf16-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__UNA-34Beagles-32K-bf16-v1.details_EleutherAI__gpt-j-6b
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__gpt-j-6b.details_declare-lab__starling-7B
Dataset Card for Evaluation run of declare-lab/starling-7B
Dataset automatically created during the evaluation run of model declare-lab/starling-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_declare-lab__starling-7B.details_SenseLLM__ReflectionCoder-DS-33B
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.details_princeton-nlp__Sheared-LLaMA-1.3B
Dataset Card for Evaluation run of princeton-nlp/Sheared-LLaMA-1.3B
Dataset Summary
Dataset automatically created during the evaluation run of model princeton-nlp/Sheared-LLaMA-1.3B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_princeton-nlp__Sheared-LLaMA-1.3B.details_migtissera__Tess-M-v1.3
Dataset Card for Evaluation run of migtissera/Tess-M-v1.3
Dataset automatically created during the evaluation run of model migtissera/Tess-M-v1.3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-M-v1.3.details_PocketDoc__Dans-TotSirocco-7b
Dataset Card for Evaluation run of PocketDoc/Dans-TotSirocco-7b
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-TotSirocco-7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-TotSirocco-7b.details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.details_togethercomputer__RedPajama-INCITE-7B-Base
Dataset Card for Evaluation run of togethercomputer/RedPajama-INCITE-7B-Base
Dataset Summary
Dataset automatically created during the evaluation run of model togethercomputer/RedPajama-INCITE-7B-Base on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_togethercomputer__RedPajama-INCITE-7B-Base.details_vilm__Quyen-Pro-Max-v0.1
Dataset Card for Evaluation run of vilm/Quyen-Pro-Max-v0.1
Dataset automatically created during the evaluation run of model vilm/Quyen-Pro-Max-v0.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_vilm__Quyen-Pro-Max-v0.1.details_meta-llama__Llama-2-70b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-70b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-70b-hf on the Open LLM Leaderboard.
The dataset is composed of 124 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-70b-hf.resume-score-details
Resume and Job Description Matching Dataset
Overview
This dataset contains 1,031 samples of resumes and job descriptions (JDs) generated and assessed using GPT-4o. The primary goal of this dataset is to evaluate the alignment between resumes and job descriptions, aiding in the study of resume relevance, skill alignment, and job fit scoring based on predefined criteria.
Dataset Composition
The dataset includes resumes matched with job descriptions, with the… See the full description on the dataset page: https://huggingface.co/datasets/netsol/resume-score-details.details_Intel__neural-chat-7b-v3-1
Dataset Card for Evaluation run of Intel/neural-chat-7b-v3-1
Dataset Summary
Dataset automatically created during the evaluation run of model Intel/neural-chat-7b-v3-1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Intel__neural-chat-7b-v3-1.details_one-man-army__una-neural-chat-v3-3-P2-OMA
Dataset Card for Evaluation run of one-man-army/una-neural-chat-v3-3-P2-OMA
Dataset automatically created during the evaluation run of model one-man-army/una-neural-chat-v3-3-P2-OMA on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__una-neural-chat-v3-3-P2-OMA.details_Azure99__blossom-v5.1-34b
Dataset Card for Evaluation run of Azure99/blossom-v5.1-34b
Dataset automatically created during the evaluation run of model Azure99/blossom-v5.1-34b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Azure99__blossom-v5.1-34b.details_AA051610__FT
Dataset Card for Evaluation run of AA051610/FT
Dataset automatically created during the evaluation run of model AA051610/FT on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AA051610__FT.details_Qwen__Qwen2-1.5B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B-Instruct.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen2-1.5B-Instruct.details_meta-llama__Llama-2-13b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-13b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-13b-hf on the Open LLM Leaderboard.
The dataset is composed of 123 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-13b-hf.details_HuggingFaceH4__zephyr-7b-beta
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset Summary
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_HuggingFaceH4__zephyr-7b-beta.details_Qwen__Qwen1.5-0.5B-Chat
Dataset Card for Evaluation run of Qwen/Qwen1.5-0.5B-Chat
Dataset automatically created during the evaluation run of model Qwen/Qwen1.5-0.5B-Chat.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen1.5-0.5B-Chat.details_EleutherAI__pythia-12b
Dataset Card for Evaluation run of EleutherAI/pythia-12b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/pythia-12b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__pythia-12b.details_01-ai__Yi-1.5-34B-Chat
Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-1.5-34B-Chat.details_openlm-research__open_llama_3b
Dataset Card for Evaluation run of openlm-research/open_llama_3b
Dataset Summary
Dataset automatically created during the evaluation run of model openlm-research/open_llama_3b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openlm-research__open_llama_3b.
