datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-Mixtral-CoT
Dataset Card for medqa-cot
Synthetically enhanced responses to the medqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.MedMCQA-Mixtral-CoT
Dataset Card for medmcqa-cot
Synthetically enhanced responses to the medmcqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedMCQA-Mixtral-CoT.mixtral-factual-QA
Mixtral Factual QA
Generate questions and answers based on context provided. We use contexts from,
maktabahalbakri.com
muftiwp.gov.my
asklegal.my
dewanbahasa-jdbp
gov.my
patriots
rootofscience
majalahsains
nasilemaktech
alhijrahnews
https://huggingface.co/datasets/open-phi/textbooks
notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/mixtral-factual
factually-wrong-qa-coding.jsonl, 31253 rows, 425 MB
factually-wrong-qa.jsonl, 1108037 rows, 10… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/mixtral-factual-QA.PubmedQA-Mixtral-CoT
Dataset Card for pubmedqa-cot
Synthetically enhanced responses to the pubmedqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the PubMedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT.Alpaca_french_mixtral
Dataset Card for Alpaca_french_mixtral
This dataset was made by reusing the french alpaca instruction with Mixtral-8x7B-Instruct to make the output open-source.
Dataset Card Contact
robinjo
clean_pubmedqa_mixtral_cot元データ: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/PubmedQA-Mixtral-CoT
データ件数: 206,962
平均トークン数: 586
最大トークン数: 1,922
合計トークン数: 121,366,170
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 532.2 MB
加工内容:
文字数によるフィルタリング:
question (質問) 列の文字数が 6,000文字を超える データを削除します。
response (応答) 列の文字数が 80,000文字を超える データを削除します。
応答 (response) の分割:
response 列を、思考プロセスを記述した「thought」部分と、最終的な結論である「answer」部分に分割します。
分割には Answer: や The answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_pubmedqa_mixtral_cot.aya_en2ro_mixtral
Dataset Card for Dataset Name
Deduplicated AYA entries translated from English to Romanian using Mixtral.
Dataset Details
Dataset Description
Curated by: https://huggingface.co/lavi13
Language(s) (NLP): Romanian
License: [More Information Needed]
Repository: [More Information Needed]
Uses
The dataset is meant be used as input data for annotation tasks. It should not be used directly for instruction tuning. It is expected to require further… See the full description on the dataset page: https://huggingface.co/datasets/lavi13/aya_en2ro_mixtral.
