CoolFace
Datasetpublic

tamarsonha/MUSE-Books-Train

MUSE-Books-Train This dataset is a simple merger of the pretraining data from the original MUSE-Books dataset. Dataset Details Dataset Sources [optional] Repository: https://huggingface.co/datasets/muse-bench/MUSE-Books Paper: https://arxiv.org/pdf/2407.06460 Dataset Creation To create this dataset, we simply started from the muse-bench dataset and selected the train subset. Then, by merging the retain1 and retain2 splits we get… See the full description on the dataset page: https://huggingface.co/datasets/tamarsonha/MUSE-Books-Train.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes61downloads
Dataset Card

MUSE-Books-Train

<!-- Provide a quick summary of the dataset. -->

This dataset is a simple merger of the pretraining data from the original MUSE-Books dataset.

Dataset Details

Dataset Sources [optional]

<!-- Provide the basic links for the dataset. -->

  • —Repository: https://huggingface.co/datasets/muse-bench/MUSE-Books
  • —Paper: https://arxiv.org/pdf/2407.06460

Dataset Creation

To create this dataset, we simply started from the muse-bench dataset and selected the train subset. Then, by merging the retain1 and retain2 splits we get the actual retain subset, and by further merging this with the original forget split, we get the full dataset used for pre-training on the specific MUSE task.