tokenized-corpus
jora_corpus1_tokenized_128kcarbon-tokenized-corpus
Dataset Summary
AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline.
Schema
The current dataset contains the following fields:
Field
Type
Description
record_id
string
Source/reference sequence identifier
start
int64
Start coordinate of the sequence interval
end
int64
End coordinate of the sequence interval
token_ids
list
Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.common-corpus-noncc-tokenizedjora_corpus1_FR_tokenized_128kPt-Corpus-Instruct-tokenized
Portuguese-Corpus Instruct (tokenized)
Dataset Summary
This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese".
For more information, see the original dataset card.
Languages
Portuguese.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.tokenized-corpus-0603
Tokenized Corpus Statistics
Shard
Total Tokens
Avg Tokens
Median Tokens
train-00000-of-00077
64,973,050
259.89
213.0
train-00001-of-00077
65,024,877
260.10
214.0
train-00002-of-00077
64,976,757
259.91
213.0
train-00003-of-00077
65,035,795
260.14
214.0
train-00004-of-00077
64,962,798
259.85
214.0
train-00005-of-00077
64,945,077
259.78
213.0
train-00006-of-00077
64,885,161
259.54
213.0
train-00007-of-00077
65,060,639
260.24
214.0
train-00008-of-0007764,905… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/tokenized-corpus-0603.
