datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
trust-game-llama-2-chat-historydetails_Norquinal__llama-2-7b-claude-chat
Dataset Card for Evaluation run of Norquinal/llama-2-7b-claude-chat
Dataset Summary
Dataset automatically created during the evaluation run of model Norquinal/llama-2-7b-claude-chat on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Norquinal__llama-2-7b-claude-chat.details_tricktreat__Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj
Dataset Card for Evaluation run of tricktreat/Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj
Dataset automatically created during the evaluation run of model tricktreat/Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tricktreat__Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj.details_Norquinal__llama-2-7b-claude-chat-rp
Dataset Card for Evaluation run of Norquinal/llama-2-7b-claude-chat-rp
Dataset Summary
Dataset automatically created during the evaluation run of model Norquinal/llama-2-7b-claude-chat-rp on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Norquinal__llama-2-7b-claude-chat-rp.DeltaSecommits_llama-2-13b-chat_tokenized_v3_vulnerabledetails_tricktreat__Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj-lora
Dataset Card for Evaluation run of tricktreat/Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj-lora
Dataset automatically created during the evaluation run of model tricktreat/Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj-lora on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tricktreat__Llama-2-7b-chat-hf-guanaco-freeze-embed-tokens-q-v-proj-lora.details_mediocredev__open-llama-3b-v2-chat
Dataset Card for Evaluation run of mediocredev/open-llama-3b-v2-chat
Dataset automatically created during the evaluation run of model mediocredev/open-llama-3b-v2-chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mediocredev__open-llama-3b-v2-chat.Taur_CoT_Analysis_Project___meta-llama__Llama-2-7b-chat-hfdetails_TheBloke__Llama-2-70B-chat-GPTQ
Dataset Card for Evaluation run of TheBloke/Llama-2-70B-chat-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Llama-2-70B-chat-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Llama-2-70B-chat-GPTQ.details_quantumaikr__llama-2-70b-fb16-orca-chat-10k
Dataset Card for Evaluation run of quantumaikr/llama-2-70b-fb16-orca-chat-10k
Dataset Summary
Dataset automatically created during the evaluation run of model quantumaikr/llama-2-70b-fb16-orca-chat-10k on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_quantumaikr__llama-2-70b-fb16-orca-chat-10k.details_wang7776__Llama-2-7b-chat-hf-10-sparsity
Dataset Card for Evaluation run of wang7776/Llama-2-7b-chat-hf-10-sparsity
Dataset Summary
Dataset automatically created during the evaluation run of model wang7776/Llama-2-7b-chat-hf-10-sparsity on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_wang7776__Llama-2-7b-chat-hf-10-sparsity.details_hamxea__Llama-2-7b-chat-hf-activity-fine-tuned-v3
Dataset Card for Evaluation run of hamxea/Llama-2-7b-chat-hf-activity-fine-tuned-v3
Dataset automatically created during the evaluation run of model hamxea/Llama-2-7b-chat-hf-activity-fine-tuned-v3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_hamxea__Llama-2-7b-chat-hf-activity-fine-tuned-v3.details_tricktreat__Llama-2-7b-chat-hf-guanaco
Dataset Card for Evaluation run of tricktreat/Llama-2-7b-chat-hf-guanaco
Dataset automatically created during the evaluation run of model tricktreat/Llama-2-7b-chat-hf-guanaco on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tricktreat__Llama-2-7b-chat-hf-guanaco.details_JCX-kcuf__Llama-2-7b-chat-hf-gpt-4-80k-base_lora
Dataset Card for Evaluation run of JCX-kcuf/Llama-2-7b-chat-hf-gpt-4-80k-base_lora
Dataset automatically created during the evaluation run of model JCX-kcuf/Llama-2-7b-chat-hf-gpt-4-80k-base_lora on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_JCX-kcuf__Llama-2-7b-chat-hf-gpt-4-80k-base_lora.details_vibhorag101__llama-2-7b-chat-hf-phr_mental_health-2048
Dataset Card for Evaluation run of vibhorag101/llama-2-7b-chat-hf-phr_mental_health-2048
Dataset Summary
Dataset automatically created during the evaluation run of model vibhorag101/llama-2-7b-chat-hf-phr_mental_health-2048 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_vibhorag101__llama-2-7b-chat-hf-phr_mental_health-2048.details_meta-llama__Llama-2-7b-chat-hf_private
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-chat-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-chat-hf.
The dataset is composed of 2 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/details_meta-llama__Llama-2-7b-chat-hf_private.details_bartowski__internlm2-chat-20b-llama
Dataset Card for Evaluation run of bartowski/internlm2-chat-20b-llama
Dataset automatically created during the evaluation run of model bartowski/internlm2-chat-20b-llama on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bartowski__internlm2-chat-20b-llama.details_BramVanroy__Llama-2-13b-chat-dutch
Dataset Card for Evaluation run of BramVanroy/Llama-2-13b-chat-dutch
Dataset Summary
Dataset automatically created during the evaluation run of model BramVanroy/Llama-2-13b-chat-dutch on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_BramVanroy__Llama-2-13b-chat-dutch.details_JCX-kcuf__Llama-2-7b-hf-llama2-chat-80k
Dataset Card for Evaluation run of JCX-kcuf/Llama-2-7b-hf-llama2-chat-80k
Dataset automatically created during the evaluation run of model JCX-kcuf/Llama-2-7b-hf-llama2-chat-80k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_JCX-kcuf__Llama-2-7b-hf-llama2-chat-80k.details_wang7776__Llama-2-7b-chat-hf-30-attention-sparsity
Dataset Card for Evaluation run of wang7776/Llama-2-7b-chat-hf-30-attention-sparsity
Dataset automatically created during the evaluation run of model wang7776/Llama-2-7b-chat-hf-30-attention-sparsity on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_wang7776__Llama-2-7b-chat-hf-30-attention-sparsity.reward-bench-Llama-2-13b-chat-hf-yes-nodetails_kfkas__Llama-2-ko-7b-Chat
Dataset Card for Evaluation run of kfkas/Llama-2-ko-7b-Chat
Dataset Summary
Dataset automatically created during the evaluation run of model kfkas/Llama-2-ko-7b-Chat on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_kfkas__Llama-2-ko-7b-Chat.meta-llama__Llama-2-7b-chat-hf-details
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-chat-hf
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-chat-hf
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Llama-2-7b-chat-hf-details.details_Lajonbot__Llama-2-7b-chat-hf-instruct-pl-lora_unload
Dataset Card for Evaluation run of Lajonbot/Llama-2-7b-chat-hf-instruct-pl-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Lajonbot/Llama-2-7b-chat-hf-instruct-pl-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Lajonbot__Llama-2-7b-chat-hf-instruct-pl-lora_unload.details_TheBloke__Llama-2-7b-Chat-AWQ
Dataset Card for Evaluation run of TheBloke/Llama-2-7b-Chat-AWQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Llama-2-7b-Chat-AWQ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Llama-2-7b-Chat-AWQ.details_lgaalves__llama-2-13b-chat-platypus
Dataset Card for Evaluation run of lgaalves/llama-2-13b-chat-platypus
Dataset Summary
Dataset automatically created during the evaluation run of model lgaalves/llama-2-13b-chat-platypus on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_lgaalves__llama-2-13b-chat-platypus.details_hamxea__Llama-2-7b-chat-hf-activity-fine-tuned-v4
Dataset Card for Evaluation run of hamxea/Llama-2-7b-chat-hf-activity-fine-tuned-v4
Dataset automatically created during the evaluation run of model hamxea/Llama-2-7b-chat-hf-activity-fine-tuned-v4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_hamxea__Llama-2-7b-chat-hf-activity-fine-tuned-v4.anthropic-hh-chosen-llama-chat-tmp2details_bartowski__internlm2-chat-7b-sft-llama
Dataset Card for Evaluation run of bartowski/internlm2-chat-7b-sft-llama
Dataset automatically created during the evaluation run of model bartowski/internlm2-chat-7b-sft-llama on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bartowski__internlm2-chat-7b-sft-llama.DeltaSecommits_llama-2-13b-chat_tokenized_v2_vulnerable
