datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresmagpie-llama-3.2-1b-instructvhab10__Llama-3.2-Instruct-3B-TIES-details
Dataset Card for Evaluation run of vhab10/Llama-3.2-Instruct-3B-TIES
Dataset automatically created during the evaluation run of model vhab10/Llama-3.2-Instruct-3B-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vhab10__Llama-3.2-Instruct-3B-TIES-details.meta-llama__Llama-3.2-3B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.2-3B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.2-3B-Instruct-details.meta-llama__Llama-3.2-1B-Instruct-details
Dataset Card for Evaluation run of meta-llama/Llama-3.2-1B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.2-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-3.2-1B-Instruct-details.agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details
Dataset Card for Evaluation run of agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
Dataset automatically created during the evaluation run of model agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details.EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-details.EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-details.open-perfectblend-llama-3.2-3b-instructqingy2024__Benchmaxx-Llama-3.2-1B-Instruct-details
Dataset Card for Evaluation run of qingy2024/Benchmaxx-Llama-3.2-1B-Instruct
Dataset automatically created during the evaluation run of model qingy2024/Benchmaxx-Llama-3.2-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/qingy2024__Benchmaxx-Llama-3.2-1B-Instruct-details.Llama-3.2-3B-Instruct-RustBusters
RustBusters Laser Cleaning QA Dataset
The RustBusters Laser Cleaning QA dataset contains 3,000 synthetic question-answer pairs designed for training a customer service assistant for RustBustersHSV, a laser cleaning and resurfacing company in Huntsville, Alabama.
Dataset Summary
This dataset consists of synthetically generated question-answer pairs designed to train a customer service assistant for a laser cleaning business. The questions cover various aspects of laser… See the full description on the dataset page: https://huggingface.co/datasets/Dudeman523/Llama-3.2-3B-Instruct-RustBusters.meditsolutions__Llama-3.2-SUN-HDIC-1B-Instruct-details
Dataset Card for Evaluation run of meditsolutions/Llama-3.2-SUN-HDIC-1B-Instruct
Dataset automatically created during the evaluation run of model meditsolutions/Llama-3.2-SUN-HDIC-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__Llama-3.2-SUN-HDIC-1B-Instruct-details.ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details.EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO-details.EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.2-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.2
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.2-details.unsloth__Llama-3.2-1B-Instruct-details
Dataset Card for Evaluation run of unsloth/Llama-3.2-1B-Instruct
Dataset automatically created during the evaluation run of model unsloth/Llama-3.2-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/unsloth__Llama-3.2-1B-Instruct-details.FuseAI__FuseChat-Llama-3.2-3B-Instruct-details
Dataset Card for Evaluation run of FuseAI/FuseChat-Llama-3.2-3B-Instruct
Dataset automatically created during the evaluation run of model FuseAI/FuseChat-Llama-3.2-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FuseAI__FuseChat-Llama-3.2-3B-Instruct-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMixBread-v0.1-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMixBread-v0.1-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMixBread-v0.1-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMixBread-v0.1-3B-details.CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-2412-details
Dataset Card for Evaluation run of CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct-2412
Dataset automatically created during the evaluation run of model CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct-2412
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-2412-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B-details.EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.3-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.3
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.3-details.win10__Llama-3.2-3B-Instruct-24-9-29-details
Dataset Card for Evaluation run of win10/Llama-3.2-3B-Instruct-24-9-29
Dataset automatically created during the evaluation run of model win10/Llama-3.2-3B-Instruct-24-9-29
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/win10__Llama-3.2-3B-Instruct-24-9-29-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B-details.insightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-details
Dataset Card for Evaluation run of insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model
Dataset automatically created during the evaluation run of model insightfactory/Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/insightfactory__Llama-3.2-3B-Instruct-unsloth-bnb-4bitlora_model-details.meditsolutions__Llama-3.2-SUN-1B-Instruct-details
Dataset Card for Evaluation run of meditsolutions/Llama-3.2-SUN-1B-Instruct
Dataset automatically created during the evaluation run of model meditsolutions/Llama-3.2-SUN-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__Llama-3.2-SUN-1B-Instruct-details.ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details
Dataset Card for Evaluation run of ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
Dataset automatically created during the evaluation run of model ModelCloud/Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ModelCloud__Llama-3.2-1B-Instruct-gptqmodel-4bit-vortex-v1-details.CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-details
Dataset Card for Evaluation run of CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct
Dataset automatically created during the evaluation run of model CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-details.Llama-3.2-3B-Instruct-nq-knowledge-probe
