CoolFace
Datasetpublic

bigcode/santacoder-token-usage

Dataset Card for "santacoder-token-usage" Token usage count per language when tokenizing the "bigcode/stack-dedup-alt-comments" dataset with the santacoder tokenizer. There are less tokens than in the tokenizer because of vocabulary mismatch between the datasets used to train the tokenizer and the ones that ended up being used to train the model. More Information needed

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes7downloads

bigcode/santacoder-token-usage · main · files are served by the source, never re-hosted here