datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-Mixtral-CoT
Dataset Card for medqa-cot
Synthetically enhanced responses to the medqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.mixtral-magicoder
Mixtral Magicoder: Source Code Is All You Need on various programming languages
We sampled programming languages from https://huggingface.co/datasets/bigcode/the-stack-dedup and pushed to https://huggingface.co/datasets/malaysia-ai/starcoderdata-sample
After that, we use Magicoder: Source Code Is All You Need on various programming languages template, we target at least 10k rows for each programming languages.
C++, 10747 rows
C#, 10193 rows
CUDA, 13843 rows
Dockerfile, 13286 rows… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-magicoder.MedMCQA-Mixtral-CoT
Dataset Card for medmcqa-cot
Synthetically enhanced responses to the medmcqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedMCQA-Mixtral-CoT.mixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.PubmedQA-Mixtral-CoT
Dataset Card for pubmedqa-cot
Synthetically enhanced responses to the pubmedqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the PubMedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT.mistralai__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.grad_school_math_instructions_fr_Mixtral
Dataset Card for grad_school_math_instructions_fr_Mixtral
This dataset was made thanks to the instruction of the vigogne's dataset but the output were generated with Mixtral-8x7B-Instruct instead of GPT3.5 to make it open-source.
Dataset Card Contact
robinjo
mixtral-malaysian-abstractive-summarization
Mixtral Malaysian Abstractive Summarization
Use https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1 to generate abstractive summarization on Malaysian dataset, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/summarization/mixtral
Example data
{'source': 'gamerbraves.com.jsonl',
'text': 'Hunter x Hunter USJ Collaboration Announced\n\n\nUniversal Studios Japan ( USJ ) has announced a collaboration with popular Shonen anime Hunter x… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-malaysian-abstractive-summarization.vicgalle__Merge-Mixtral-Prometheus-8x7B-details
Dataset Card for Evaluation run of vicgalle/Merge-Mixtral-Prometheus-8x7B
Dataset automatically created during the evaluation run of model vicgalle/Merge-Mixtral-Prometheus-8x7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Merge-Mixtral-Prometheus-8x7B-details.grad_school_math_instructions_fr_Mixtral
Dataset Card for grad_school_math_instructions_fr_Mixtral
This dataset was made thanks to the instruction of the vigogne's dataset but the output were generated with Mixtral-8x7B-Instruct instead of GPT3.5 to make it open-source.
Dataset Card Contact
robinjo
mistralai__Mixtral-8x7B-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x7B-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x7B-v0.1
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x7B-v0.1-details.mistralai__Mixtral-8x7B-Instruct-v0.1-details
Dataset Card for Evaluation run of mistralai/Mixtral-8x7B-Instruct-v0.1
Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x7B-Instruct-v0.1
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x7B-Instruct-v0.1-details.LeroyDyer__Mixtral_AI_SwahiliTron_7b-details
Dataset Card for Evaluation run of LeroyDyer/Mixtral_AI_SwahiliTron_7b
Dataset automatically created during the evaluation run of model LeroyDyer/Mixtral_AI_SwahiliTron_7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__Mixtral_AI_SwahiliTron_7b-details.Alpaca_french_mixtral
Dataset Card for Alpaca_french_mixtral
This dataset was made by reusing the french alpaca instruction with Mixtral-8x7B-Instruct to make the output open-source.
Dataset Card Contact
robinjo
NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details.mistral-community__mixtral-8x22B-v0.3-details
Dataset Card for Evaluation run of mistral-community/mixtral-8x22B-v0.3
Dataset automatically created during the evaluation run of model mistral-community/mixtral-8x22B-v0.3
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistral-community__mixtral-8x22B-v0.3-details.clean_pubmedqa_mixtral_cot元データ: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/PubmedQA-Mixtral-CoT
データ件数: 206,962
平均トークン数: 586
最大トークン数: 1,922
合計トークン数: 121,366,170
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 532.2 MB
加工内容:
文字数によるフィルタリング:
question (質問) 列の文字数が 6,000文字を超える データを削除します。
response (応答) 列の文字数が 80,000文字を超える データを削除します。
応答 (response) の分割:
response 列を、思考プロセスを記述した「thought」部分と、最終的な結論である「answer」部分に分割します。
分割には Answer: や The answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_pubmedqa_mixtral_cot.NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT
The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-SFT-details.abacusai__Smaug-Mixtral-v0.1-details
Dataset Card for Evaluation run of abacusai/Smaug-Mixtral-v0.1
Dataset automatically created during the evaluation run of model abacusai/Smaug-Mixtral-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Smaug-Mixtral-v0.1-details.mistral-community__Mixtral-8x22B-v0.1-details
Dataset Card for Evaluation run of mistral-community/Mixtral-8x22B-v0.1
Dataset automatically created during the evaluation run of model mistral-community/Mixtral-8x22B-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistral-community__Mixtral-8x22B-v0.1-details.ai-arxiv2-ragas-mixtralcloudyu__Mixtral_34Bx2_MoE_60B-details
Dataset Card for Evaluation run of cloudyu/Mixtral_34Bx2_MoE_60B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_34Bx2_MoE_60B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_34Bx2_MoE_60B-details.OpenBuddy__openbuddy-mixtral-7bx8-v18.1-32k-details
Dataset Card for Evaluation run of OpenBuddy/openbuddy-mixtral-7bx8-v18.1-32k
Dataset automatically created during the evaluation run of model OpenBuddy/openbuddy-mixtral-7bx8-v18.1-32k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenBuddy__openbuddy-mixtral-7bx8-v18.1-32k-details.research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.mixtral-distillation-tools-dataCreativeCommons-RAG-QA-Mixtral8x22b
以下のデータ源からランダムに抽出した日本語のテキストをもとに、RAG形式のQ&Aを自動生成したものです。
Wikibooks
Wikipedia
判例データ
instruction datasetとしてではなく、事前学習での利用を想定しています(質疑応答をするための訓練)。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
aya_en2ro_mixtral
Dataset Card for Dataset Name
Deduplicated AYA entries translated from English to Romanian using Mixtral.
Dataset Details
Dataset Description
Curated by: https://huggingface.co/lavi13
Language(s) (NLP): Romanian
License: [More Information Needed]
Repository: [More Information Needed]
Uses
The dataset is meant be used as input data for annotation tasks. It should not be used directly for instruction tuning. It is expected to require further… See the full description on the dataset page: https://huggingface.co/datasets/lavi13/aya_en2ro_mixtral.conceptnet_UsedFor_en_en_mixtral_finetuneThe purpose of this dataset is to be used for a fine tuning on Mixtral, it contains all the 'UsedFor' relationships (english to english) present in ConceptNet 5.7.0 in the form of aggregated instructions, i.e. for any arg1, arg2_list is the list of all arg2 found to be in a UsedFor relationship with arg1 (arg1 --UsedFor--> arg2)
The instruction is written in the following format: <s> [INST] instruction [/INST] answer </s>
cloudyu__Mixtral_11Bx2_MoE_19B-details
Dataset Card for Evaluation run of cloudyu/Mixtral_11Bx2_MoE_19B
Dataset automatically created during the evaluation run of model cloudyu/Mixtral_11Bx2_MoE_19B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cloudyu__Mixtral_11Bx2_MoE_19B-details.doom-mixtral-text
