Rubin-Wei/enwiki-dec2021-preprocessed-mistral
Dataset Description This dataset is a preprocessed version of the English Wikipedia snapshot from December 2021, processed using the preprocess_dataset.py script provided in the repository below. Paper: MLP Memory: A Retriever-Pretrained Memory for Large Language Models GitHub: https://github.com/Rubin-Wei/MLPMemory Dataset Source: English Wikipedia (December 2021) Tokenizer: Mistral-7B-v0.3 Two key preprocessing parameters used are: block_size: 2048 stride: 1024… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/enwiki-dec2021-preprocessed-mistral.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face