ttj/dataset-ironing-smollm2-135m-ppl
base_data_smolm2_135m_ppl This dataset is a nanochat pretraining parquet dataset annotated with offline per-example perplexity from HuggingFaceTB/SmolLM2-135M. Columns text: original training text. perplexity: per-example perplexity from the fixed reference model. perplexity_num_tokens: number of reference-token prediction targets used. perplexity_nll: mean negative log likelihood before exponentiation. perplexity_reference_model: reference model id used for the… See the full description on the dataset page: https://huggingface.co/datasets/ttj/dataset-ironing-smollm2-135m-ppl.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face