CoolFace
Datasetpublic

bigcode/santacoder-token-usage

Dataset Card for "santacoder-token-usage" Token usage count per language when tokenizing the "bigcode/stack-dedup-alt-comments" dataset with the santacoder tokenizer. There are less tokens than in the tokenizer because of vocabulary mismatch between the datasets used to train the tokenizer and the ones that ended up being used to train the model. More Information needed

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes7downloads
4 commits on main
7e0fd004y ago

Update README.md

cakiki
0a0050d4y ago

Upload README.md with huggingface_hub

cakiki
1999e104y ago

Upload data/train-00000-of-00001-02a8910126c18b3a.parquet with huggingface_hub

cakiki
873af344y ago

initial commit

christopher