mnlp
Datasets
All datasets matching “mnlp”MNLP_M3_mcqa_datasetMNLP_M2_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets:
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the stem subsets
mmlu we only use the stem subsets
ai2_arc
ScienceQA
math_qa… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M2_mcqa_dataset.MNLP_M3_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets (and the test and validation splits are only used for testing not for training):
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M3_mcqa_dataset.mmlu-pro-augmentationMNLP_MCQA_datasetThis MCQA dataset (of only single answer) contains a mixture of train, validation and test from this datasets (test and validation are only used for testing not for training):
mmlu auxiliary train Only the stem subset is used
mmlu Only the stem subset is used
mmlu 10 choices auxiliary train stem
ai2_arc
ScienceQA
math_qa
openbook_qa
sciq
medmcqa A 32,000 random subset (seed 42)
Instruction-finetuning-mixture-mnlpDataset created using the Tulu3-sft-mixture
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
