datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MNLP_M3_mcqa_dataset_openbookqa_cotMNLP_M3_mcqa_dataset_openbookqa_origMNLP_M3_quantized_dataset
Enhanced MCQA Test Dataset for Comprehensive Model Evaluation
This dataset contains 400 carefully selected test samples from MetaMathQA, AQuA-RAT, OpenBookQA, and SciQ datasets, designed for comprehensive MCQA (Multiple Choice Question Answering) model evaluation and quantization testing across multiple domains.
Dataset Overview
Total Samples: 400
MetaMathQA Samples: 100 (mathematical problems)
AQuA-RAT Samples: 100 (algebraic word problems)
OpenBookQA Samples: 100… See the full description on the dataset page: https://huggingface.co/datasets/AlirezaAbdollahpoor/MNLP_M3_quantized_dataset.MNLP_M3_RAG_corpusMNLP_M3_quantized_datasetMNLP_M3_extra_data_bioMNLP_M3_extra_data_sciqMNLP_M3_rag_documents_300tokMNLP_M3_rag_documentsMNLP_M3_dpo_datasetMNLP_M3_extra_data_medicineMNLP_M3_mcqa_datasetMNLP_M3_dpo_datasetMNLP_M3_dpo_datasetThis dataset was created for training and evaluating a DPO-based language model in the context of STEM questions. It supports both instruction tuning and preference-based fine-tuning using the DPO framework. The dataset was developed for the CS-552 course Modern Natural Language Processing.
Dataset structure :
The default subset contains the DPO data (preference pairs). This data comes from Milestone 1 (pref pairs collected by students) or from different dpo datasets available on… See the full description on the dataset page: https://huggingface.co/datasets/madhueb/MNLP_M3_dpo_dataset.
