datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rpj-v2-sample-mixtralMedQA-Mixtral-CoT
Dataset Card for medqa-cot
Synthetically enhanced responses to the medqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.mixtral-chunkstinygsm_mixtral_12Mdetails_VAGOsolutions__SauerkrautLM-Mixtral-8x7B-Instruct
Dataset Card for Evaluation run of VAGOsolutions/SauerkrautLM-Mixtral-8x7B-Instruct
Dataset automatically created during the evaluation run of model VAGOsolutions/SauerkrautLM-Mixtral-8x7B-Instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_VAGOsolutions__SauerkrautLM-Mixtral-8x7B-Instruct.synthetic_zeroshot_mixtral_v0.1details_hfl__chinese-mixtral
Dataset Card for Evaluation run of hfl/chinese-mixtral
Dataset automatically created during the evaluation run of model hfl/chinese-mixtral on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_hfl__chinese-mixtral.mixtral-magicoder
Mixtral Magicoder: Source Code Is All You Need on various programming languages
We sampled programming languages from https://huggingface.co/datasets/bigcode/the-stack-dedup and pushed to https://huggingface.co/datasets/malaysia-ai/starcoderdata-sample
After that, we use Magicoder: Source Code Is All You Need on various programming languages template, we target at least 10k rows for each programming languages.
C++, 10747 rows
C#, 10193 rows
CUDA, 13843 rows
Dockerfile, 13286 rows… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-magicoder.3_4_fusechat_v1_openchat-3.5_nh2-mixtral-8x7b-dpo_nh2-solar-10.7b_representation1_4_fusechat_v1_openchat-3.5_nh2-mixtral-8x7b-dpo_nh2-solar-10.7b_representationdetails_mistralai__Mixtral-8x7B-v0.1
Dataset Card for Evaluation run of mistralai/Mixtral-8x7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x7B-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mistralai__Mixtral-8x7B-v0.1.details_LeroyDyer__Mixtral_AI_Cyber_4.0
Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_Cyber_4.0
Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_Cyber_4.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_LeroyDyer__Mixtral_AI_Cyber_4.0.MedMCQA-Mixtral-CoT
Dataset Card for medmcqa-cot
Synthetically enhanced responses to the medmcqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedMCQA-Mixtral-CoT.details_vistagi__Mixtral-8x7b-v0.1-sft
Dataset Card for Evaluation run of vistagi/Mixtral-8x7b-v0.1-sft
Dataset automatically created during the evaluation run of model vistagi/Mixtral-8x7b-v0.1-sft on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_vistagi__Mixtral-8x7b-v0.1-sft.Mixtral-8x7B-v0.1-activationsdetails_Brillibits__Instruct_Mixtral-8x7B-v0.1_Dolly15K
Dataset Card for Evaluation run of Brillibits/Instruct_Mixtral-8x7B-v0.1_Dolly15K
Dataset automatically created during the evaluation run of model Brillibits/Instruct_Mixtral-8x7B-v0.1_Dolly15K on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Brillibits__Instruct_Mixtral-8x7B-v0.1_Dolly15K.alpaca-french-mixtral
License & Attribution
MTEB-format derivative of AIffl/Alpaca_french_mixtral (French Alpaca, Mixtral-translated). Query = instruction; corpus = answer. Deterministically subsampled to ~10k. Licensed under Apache-2.0 (same as source).
3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationdetails_Swisslex__Mixtral-Orca-v0.1
Dataset Card for Evaluation run of Swisslex/Mixtral-Orca-v0.1
Dataset automatically created during the evaluation run of model Swisslex/Mixtral-Orca-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Swisslex__Mixtral-Orca-v0.1.details_Swisslex__Mixtral-8x7b-DPO-v0.2
Dataset Card for Evaluation run of Swisslex/Mixtral-8x7b-DPO-v0.2
Dataset automatically created during the evaluation run of model Swisslex/Mixtral-8x7b-DPO-v0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Swisslex__Mixtral-8x7b-DPO-v0.2.mixtraltoken_fineweb_edu_mini_combinedv3-mix-mixtralmixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.mixtral-malaysian-general-qa
Mixtral Malaysian Chat
Simulate conversation between a user and an assistant on various topics. Generated using Mixtral Instructions.
Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/chatbot/mixtral-malaysian-chat
Multi-turn Bad things
Multiturn of the user is saying bad things to the assistant.
mixtral-conversation-badthings.jsonl, 57798 rows, 163 MB.
Example data
[{'role': 'user',
'content': "Hey bot, you're really dumb."… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-malaysian-general-qa.details_cloudyu__Mixtral_7Bx2_MoE_13B
Dataset Card for Evaluation run of cloudyu/Mixtral_7Bx2_MoE_13B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_7Bx2_MoE_13B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cloudyu__Mixtral_7Bx2_MoE_13B.details_cloudyu__Mixtral_7Bx5_MoE_30B
Dataset Card for Evaluation run of cloudyu/Mixtral_7Bx5_MoE_30B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_7Bx5_MoE_30B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cloudyu__Mixtral_7Bx5_MoE_30B.details_cloudyu__Mixtral_34Bx2_MoE_60B
Dataset Card for Evaluation run of cloudyu/Mixtral_34Bx2_MoE_60B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_34Bx2_MoE_60B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cloudyu__Mixtral_34Bx2_MoE_60B.details_Sao10K__Sensualize-Mixtral-bf16
Dataset Card for Evaluation run of Sao10K/Sensualize-Mixtral-bf16
Dataset automatically created during the evaluation run of model Sao10K/Sensualize-Mixtral-bf16 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Sao10K__Sensualize-Mixtral-bf16.details_Open-Orca__Mixtral-SlimOrca-8x7B
Dataset Card for Evaluation run of Open-Orca/Mixtral-SlimOrca-8x7B
Dataset automatically created during the evaluation run of model Open-Orca/Mixtral-SlimOrca-8x7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Open-Orca__Mixtral-SlimOrca-8x7B.details_chargoddard__MixtralRPChat-ZLoss
Dataset Card for Evaluation run of chargoddard/MixtralRPChat-ZLoss
Dataset automatically created during the evaluation run of model chargoddard/MixtralRPChat-ZLoss on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_chargoddard__MixtralRPChat-ZLoss.
