CoolFace
Datasetpublic

sycmucmu/prolong-smollm2-validation

ProLong SmolLM2 validation Two validation sets for evaluating DCLM-trained language models, retokenized from the code, books, and textbooks subsets of princeton-nlp/prolong-data-64K. Folder Context length Sequences Usable tokens Stored tokens Trailing EOS filler (unused) 4k/ 4,096 24,414 99,999,744 100,000,000 256 32k/ 32,768 3,051 99,975,168 100,000,000 24,832 Each folder contains prolong_val_100m.bin, per-sequence source metadata in prolong_val_100m.json, and… See the full description on the dataset page: https://huggingface.co/datasets/sycmucmu/prolong-smollm2-validation.

sourceHugging Faceupdated 10d agoView on Hugging Face
0likes330downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
sycmucmu/prolong-smollm2-validation · CoolFace