emozilla/dolma-v1_7-3B-tokenized-llama3-nanoset
Tokenized (Llama 3) verison of NousResearch/dolma-v1_7-3B as a Nanotron dataset. Can also be used directly with numpy, for example import numpy as np dataset_buffer_mmap = np.memmap("dolma-v1_7-3B-nanoset-l3_input_ids.npy", mode="r", order="C", dtype=np.int32) dataset_buffer = memoryview(dataset_buffer_mmap) dataset_number_of_tokens = int(len(dataset_buffer))
118
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face