datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MemoryDecoder-at-Scale-domain-data
MemoryDecoder at Scale Domain Data
This repository contains the domain-specific continued-pretraining (CPT) data,
the tokenized and preprocessed datasets, and the aligned KNN distributions used
by MemoryDecoder at Scale.
Links
Project Page: Memory Decoder at Scale
GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale
Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.MemoryDecoder-domain-data
Dataset Description
This dataset contains the test splits used to evaluate the Memory Decoder model across three specialized domains: biomedical, legal, and finance.
The test data was randomly sampled from publicly available datasets to assess the model's performance in domain-specific language understanding.
GitHub: https://github.com/LUMIA-Group/MemoryDecoder
Dataset Sources
The test data is sampled randomly from the following source datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Clover-Hill/MemoryDecoder-domain-data.
