datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MNLP_M3_mcqa_datasetMNLP_M2_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets:
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the stem subsets
mmlu we only use the stem subsets
ai2_arc
ScienceQA
math_qa… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M2_mcqa_dataset.MNLP_M3_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets (and the test and validation splits are only used for testing not for training):
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M3_mcqa_dataset.mmlu-pro-augmentationMNLP_MCQA_datasetThis MCQA dataset (of only single answer) contains a mixture of train, validation and test from this datasets (test and validation are only used for testing not for training):
mmlu auxiliary train Only the stem subset is used
mmlu Only the stem subset is used
mmlu 10 choices auxiliary train stem
ai2_arc
ScienceQA
math_qa
openbook_qa
sciq
medmcqa A 32,000 random subset (seed 42)
Instruction-finetuning-mixture-mnlpDataset created using the Tulu3-sft-mixture
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
MNLP_M3_dpo_datasetMNLP_M3_rag_datasetMNLP_M3_mcqa_datasetMNLP_M3_mcqa_dataset
Tulu 3 SFT Mixture (Sampled)
This dataset is a sampled and filtered subset of the allenai/tulu-3-sft-mixture, curated and rebalanced for structured instruction fine-tuning. The goal is to support research and model development in math reasoning, coding, knowledge recall, instruction following (IF), and conversational alignment, while explicitly excluding safety, multilingual, and certain task-specific sources.
📦 Dataset Structure
Source: Filtered from… See the full description on the dataset page: https://huggingface.co/datasets/vanek-epfl/MNLP_M3_mcqa_dataset.MNLP_M2_dpo_datasetMNLP_M3_mcqa_datasetMNLP_M2_mcqa_datasetMNLP_M3_rag_datasetMNLP_MCQA_dataset_2This MCQA dataset (of only single answer) contains a mixture of train, validation and test from this datasets (test and validation are only used for testing not for training):
mmlu auxiliary train Only the stem subset is used
mmlu Only the stem subset is used
ai2_arc
ScienceQA
math_qa
openbook_qa
sciq
medmcqa
Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
MNLP_intstruction_tuningMNLP_M3_quantized_datasethw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.MNLP_M3_quantized_datasetMNLP_M2_quantized_dataset1MNLP_M3_mcqa_dataset_openbookqa_cotMNLP_M2_quantized_dataset
Dataset Card for MNLP_M2_sft_dataset
Dataset Description
A unified STEM instruction-following dataset comprising 240,500 examples drawn from six existing benchmarks: SciQ, Deepmind Code Contests, TIGER-Lab MathInstruct, TULU Algebra, TULU Code, and Facebook Natural Reasoning. Each example is formatted as a chat-style message pair for supervised fine-tuning of instruction-following models.
Curated by: Sarra Chabane
Shared by: GingerBled (https://huggingface.co/GingerBled)… See the full description on the dataset page: https://huggingface.co/datasets/arthurrpp/MNLP_M2_quantized_dataset.MNLP_M3_mcqa_dataset_qasc_cotMNLP_M2_documentsMNLP_M2_quantized_datasetMNLP_M2_quantized_datasetsuper_gpqa_mnlp_m3MNLP_M3_generator_trainingInstruction-finetuning-mixture-mnlp-only-english-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
The dataset is only with english language
