bigcode/santacoder-token-usage
Dataset Card for "santacoder-token-usage" Token usage count per language when tokenizing the "bigcode/stack-dedup-alt-comments" dataset with the santacoder tokenizer. There are less tokens than in the tokenizer because of vocabulary mismatch between the datasets used to train the tokenizer and the ones that ended up being used to train the model. More Information needed
07
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face