datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresllama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.llama-3.1-awesome-chatgpt-prompts
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.Yuma42__Llama3.1-SuperHawk-8B-details
Dataset Card for Evaluation run of Yuma42/Llama3.1-SuperHawk-8B
Dataset automatically created during the evaluation run of model Yuma42/Llama3.1-SuperHawk-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Yuma42__Llama3.1-SuperHawk-8B-details.netcat420__MFANN-llama3.1-abliterated-SLERP-v3-details
Dataset Card for Evaluation run of netcat420/MFANN-llama3.1-abliterated-SLERP-v3
Dataset automatically created during the evaluation run of model netcat420/MFANN-llama3.1-abliterated-SLERP-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/netcat420__MFANN-llama3.1-abliterated-SLERP-v3-details.MaziyarPanahi__calme-2.3-llama3.1-70b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3.1-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3.1-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.3-llama3.1-70b-details.sequelbox__Llama3.1-8B-PlumCode-details
Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumCode
Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumCode
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumCode-details.nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details
Dataset Card for Evaluation run of nbeerbower/Llama3.1-Gutenberg-Doppel-70B
Dataset automatically created during the evaluation run of model nbeerbower/Llama3.1-Gutenberg-Doppel-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__Llama3.1-Gutenberg-Doppel-70B-details.nbeerbower__llama3.1-kartoffeldes-70B-details
Dataset Card for Evaluation run of nbeerbower/llama3.1-kartoffeldes-70B
Dataset automatically created during the evaluation run of model nbeerbower/llama3.1-kartoffeldes-70B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nbeerbower__llama3.1-kartoffeldes-70B-details.agentlans__Llama3.1-Daredevilish-Instruct-details
Dataset Card for Evaluation run of agentlans/Llama3.1-Daredevilish-Instruct
Dataset automatically created during the evaluation run of model agentlans/Llama3.1-Daredevilish-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama3.1-Daredevilish-Instruct-details.EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details.EpistemeAI__Alpaca-Llama3.1-8B-details
Dataset Card for Evaluation run of EpistemeAI/Alpaca-Llama3.1-8B
Dataset automatically created during the evaluation run of model EpistemeAI/Alpaca-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Alpaca-Llama3.1-8B-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details.EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details.EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.03-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.03-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.03-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.03-8B-Philos-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.04-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.04-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.04-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.04-8B-Philos-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-C-R1-KTO-Reflection-details.llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.EpistemeAI2__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.01-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.01-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.01-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.01-8B-Philos-details.llama-3.1-medprm-reward-raw-training-setnecva__IE-cont-Llama3.1-8B-details
Dataset Card for Evaluation run of necva/IE-cont-Llama3.1-8B
Dataset automatically created during the evaluation run of model necva/IE-cont-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/necva__IE-cont-Llama3.1-8B-details.agentlans__Llama3.1-LexiHermes-SuperStorm-details
Dataset Card for Evaluation run of agentlans/Llama3.1-LexiHermes-SuperStorm
Dataset automatically created during the evaluation run of model agentlans/Llama3.1-LexiHermes-SuperStorm
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama3.1-LexiHermes-SuperStorm-details.Yuma42__Llama3.1-IgneousIguana-8B-details
Dataset Card for Evaluation run of Yuma42/Llama3.1-IgneousIguana-8B
Dataset automatically created during the evaluation run of model Yuma42/Llama3.1-IgneousIguana-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Yuma42__Llama3.1-IgneousIguana-8B-details.cognitivecomputations__Dolphin3.0-Llama3.1-8B-details
Dataset Card for Evaluation run of cognitivecomputations/Dolphin3.0-Llama3.1-8B
Dataset automatically created during the evaluation run of model cognitivecomputations/Dolphin3.0-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__Dolphin3.0-Llama3.1-8B-details.Nexesenex__Dolphin3.0-Llama3.1-1B-abliterated-details
Dataset Card for Evaluation run of Nexesenex/Dolphin3.0-Llama3.1-1B-abliterated
Dataset automatically created during the evaluation run of model Nexesenex/Dolphin3.0-Llama3.1-1B-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Dolphin3.0-Llama3.1-1B-abliterated-details.nvidia__OpenMath2-Llama3.1-8B-details
Dataset Card for Evaluation run of nvidia/OpenMath2-Llama3.1-8B
Dataset automatically created during the evaluation run of model nvidia/OpenMath2-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nvidia__OpenMath2-Llama3.1-8B-details.
