datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoA_Long_HumanQA
MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression
This is the dataset used by the automatic sparse attention compression method MoA.
It enhances the calibration dataset by integrating long-range dependencies and model alignment.
MoA utilizes long-contextual datasets, which include question-answer pairs heavily dependent on long-range content.
The question-answer pairs are written by human in this dataset repository. Large language Models (LLMs) should… See the full description on the dataset page: https://huggingface.co/datasets/nics-efc/MoA_Long_HumanQA.MoA_Long_RetrievalPersian-moadl
🎯 Persian Math Questions Dataset for SFT
📝 Description
This dataset contains Persian questions primarily focused on mathematical concepts, designed for Supervised Fine-Tuning (SFT) of Language Models.
🔍 Features
High-quality Persian questions
Detailed subtopic categorization
Focused on mathematical concepts
Tokens count for each conversation
🚀 Coming Soon
Detailed answers for each question
Additional topics beyond mathematics
Enhanced… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/Persian-moadl.
