datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m2d2-wiki-decon
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/claran/m2d2-wiki-decon.clara-stage2-data
Clara Stage 2 Training Data
Training data for Clara Stage 2 (Compression Instruction Tuning).
Dataset Description
This dataset contains high-quality QA pairs with single documents for training Clara's decoder adapter to generate answers from compressed document representations.
Data Format
Each record contains:
question: The query/question
answer: Gold answer
docs: List containing 1 document
meta: Source description
metadata: Additional metadata (repo, scope… See the full description on the dataset page: https://huggingface.co/datasets/dl3239491/clara-stage2-data.seed-pretrain-decon
Dataset Card for Dataset Name
Pre-training corpus for seed models in "Scalable Data Ablation Approximations for Language Models through Modular Training and Merging", to be presented at EMNLP 2024.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/claran/seed-pretrain-decon.
