CoolFace
6 results

tokenized-corpus

jinofy-corp /jora_corpus1_tokenized_128ktabularn<1K4 likes42k downloads2mo agoHugging FaceAINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline. Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes865 downloads23d agoHugging Faceynnekuw /common-corpus-noncc-tokenized10M<n<100M0 likes793 downloads3mo agoHugging Facejinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes614 downloads2mo agoHugging FacenicholasKluge /Pt-Corpus-Instruct-tokenized Portuguese-Corpus Instruct (tokenized) Dataset Summary This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese". For more information, see the original dataset card. Languages Portuguese. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.text-generation1M<n<10M0 likes584 downloads1y agoHugging FaceBa2han /tokenized-corpus-0603 Tokenized Corpus Statistics Shard Total Tokens Avg Tokens Median Tokens train-00000-of-00077 64,973,050 259.89 213.0 train-00001-of-00077 65,024,877 260.10 214.0 train-00002-of-00077 64,976,757 259.91 213.0 train-00003-of-00077 65,035,795 260.14 214.0 train-00004-of-00077 64,962,798 259.85 214.0 train-00005-of-00077 64,945,077 259.78 213.0 train-00006-of-00077 64,885,161 259.54 213.0 train-00007-of-00077 65,060,639 260.24 214.0 train-00008-of-0007764,905… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/tokenized-corpus-0603.10M<n<100M0 likes484 downloads7mo agoHugging Face