Rubin-Wei/enwiki-dec2021-preprocessed-mistral
Dataset Description This dataset is a preprocessed version of the English Wikipedia snapshot from December 2021, processed using the preprocess_dataset.py script provided in the repository below. Paper: MLP Memory: A Retriever-Pretrained Memory for Large Language Models GitHub: https://github.com/Rubin-Wei/MLPMemory Dataset Source: English Wikipedia (December 2021) Tokenizer: Mistral-7B-v0.3 Two key preprocessing parameters used are: block_size: 2048 stride: 1024… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/enwiki-dec2021-preprocessed-mistral.
This repository belongs to Rubin-Wei on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
