jiluoaaron/CrossFit-and-UnifiedQA
Dataset Card for CrossFit-and-UnifiedQA Dataset Summary CrossFit is a benchmark dedicated to evaluating cross-task generalization in few-shot NLP learning. It establishes a standardized evaluation paradigm and integrates 160 diverse few-shot tasks into a unified text-to-text format via NLP Few-shot Gym, facilitating reliable assessment of model generalization. UnifiedQA aims to break format boundaries in QA research by unifying over 20 datasets across four… See the full description on the dataset page: https://huggingface.co/datasets/jiluoaaron/CrossFit-and-UnifiedQA.
Dataset Card for CrossFit-and-UnifiedQA
Dataset Description
- Homepage:
- Repository: https://github.com/AaronJi/MeGan
- Paper: https://arxiv.org/abs/2605.01973
- Leaderboard:
- Point of Contact:
Dataset Summary
CrossFit is a benchmark dedicated to evaluating cross-task generalization in few-shot NLP learning. It establishes a standardized evaluation paradigm and integrates 160 diverse few-shot tasks into a unified text-to-text format via NLP Few-shot Gym, facilitating reliable assessment of model generalization. UnifiedQA aims to break format boundaries in QA research by unifying over 20 datasets across four mainstream QA formats. It enables consistent evaluation of QA models’ generalization ability without task-specific customization.
Supported Tasks
The dataset includes 7 meta-task settings:
- hr->lr
- class->class
- non-class->class
- qa->qa
- non-qa->qa
- non-NLI->NLI
- non-paraphrase->paraphrase
Detailed information can be found in the original paper: MetaICL: Learning to learn in context, NAACL 2022
Languages
English
Dataset Structure
The data has a standard messages (sharegpt) formart.
Data Instances
{ "messages": [ { "role": "user", "content": "sentence 1: An introduction to atoms and elements, compounds, atomic structure and bonding, the molecule and chemical reactions. sentence 2: Replace another in a molecule happens to atoms during a substitution reaction.", "instruction": "which of the following answers is the most appropriate: [entailment; neutral; ]" }, { "role": "assistant", "content": "neutral" } ], "task": "scitail", "context": "", "task_description": "Examine the semantic and logical consistency of a proposed hypothesis with the given premise. Your goal is to determine the strongest epistemic relation between the two statements." }
Data Fields
- role: user/assistant
- content: the question
- instruction: the instruction to answer the question (maybe empty)
- context: the context to answer the question (maybe empty)
- task: the detailed task
- task_description: the description of the specific tasks, as the textual input of the hypernetwork in MeGan (ICML2026).
Data Splits
- train/trainfix: the training split (fix: with length filtering)
- test/testfix: the test split (fix: with length filtering)
- unseentest/unseentestfix: the test split with only tasks unseen in the training (fix: with length filtering)
