ttj/dataset-ironing-smollm2-135m-ppl
base_data_smolm2_135m_ppl This dataset is a nanochat pretraining parquet dataset annotated with offline per-example perplexity from HuggingFaceTB/SmolLM2-135M. Columns text: original training text. perplexity: per-example perplexity from the fixed reference model. perplexity_num_tokens: number of reference-token prediction targets used. perplexity_nll: mean negative log likelihood before exponentiation. perplexity_reference_model: reference model id used for the… See the full description on the dataset page: https://huggingface.co/datasets/ttj/dataset-ironing-smollm2-135m-ppl.
This repository belongs to ttj on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
