datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jora_corpus1_tokenized_128kcarbon-tokenized-corpus
Dataset Summary
AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline.
Schema
The current dataset contains the following fields:
Field
Type
Description
record_id
string
Source/reference sequence identifier
start
int64
Start coordinate of the sequence interval
end
int64
End coordinate of the sequence interval
token_ids
list
Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.common-corpus-noncc-tokenizedjora_corpus1_FR_tokenized_128kPt-Corpus-Instruct-tokenized
Portuguese-Corpus Instruct (tokenized)
Dataset Summary
This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese".
For more information, see the original dataset card.
Languages
Portuguese.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.tokenized-corpus-0603
Tokenized Corpus Statistics
Shard
Total Tokens
Avg Tokens
Median Tokens
train-00000-of-00077
64,973,050
259.89
213.0
train-00001-of-00077
65,024,877
260.10
214.0
train-00002-of-00077
64,976,757
259.91
213.0
train-00003-of-00077
65,035,795
260.14
214.0
train-00004-of-00077
64,962,798
259.85
214.0
train-00005-of-00077
64,945,077
259.78
213.0
train-00006-of-00077
64,885,161
259.54
213.0
train-00007-of-00077
65,060,639
260.24
214.0
train-00008-of-0007764,905… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/tokenized-corpus-0603.turkish_corpus_tokenized
Dataset Card for "turkish_corpus_tokenized"
More Information needed
sindhi-gold-corpus-mlm-tokenized-for-bertllama2-german-corpus-tokenized-llama-chunk-4096
Dataset Card for "llama2-german-corpus-tokenized-llama-chunk-4096"
More Information needed
Pt-Corpus-tokenized
Portuguese-Corpus (tokenized)
Dataset Summary
This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus dataset. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese".
For more information, see the original dataset card.
Languages
Portuguese.
Dataset Structure
Data Instances
The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-tokenized.long_corpus-0209_tokenized
long_corpus-0209_tokenized
Exact-deduplicated, tokenized, and shuffled pretraining mix.
Processing
Tokenizer: /workspace/good_tokenizer (vocab_size=60800)
Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3)
bos_eos = True: every example is [BOS] + content + [EOS]
min_tokens = 15 (including BOS/EOS)
max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos)
Dedup: exact match on stripped UTF-8 text (blake2s-128)
Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.tokenized_large_corpus_v2
Dataset Card for "tokenized_large_corpus_v2"
More Information needed
vietnamese-news-corpus-tokenizedsindhi-gold-corpus-mlm-tokenized-mberth-corpus-tokenized-qwen2.5-ctx2048large_spanish_corpus_ds_tokenized_and_gropuped
Dataset Card for "large_spanish_corpus_ds_tokenized_and_gropuped"
More Information needed
magnetar-tokenized-corpuscorpus_clean_tokenized
Kazakh Tokenized Corpus (2048 blocks)
Pre-tokenized Kazakh corpus ready for language model training. Built from Abzalbek89/corpus_clean using Abzalbek89/kk-tokenizer-bpe-32k.
Dataset Summary
Metric
Value
Train blocks
236,981
Validation blocks
12,473
Total blocks
249,454
Block size
2,048 tokens
Total tokens
~0.51B (510M)
Tokenizer
ByteLevel BPE, 32K vocab
Val ratio
5%
Pipeline
Source: Abzalbek89/corpus_clean — 1,502,583 cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Abzalbek89/corpus_clean_tokenized.common-corpus-noncc-500b-tokenizedvngrs-web-corpus-500k-kumru_tokenizer-tokenizedtelugu-colloquial-corpus-tokenized
Telugu Colloquial Corpus (Tokenized)
This dataset is a tokenized version of the Telugu Colloquial Corpus (TeCC). It contains examples of informal, everyday Telugu language, including slang, regional variations, and conversational patterns.
Dataset Details
Language: Telugu (te)
Tokenization: Tokenized using the bert-base-multilingual-cased tokenizer from the transformers library.
Source: [Describe where the original data came from – e.g., collected from online forums… See the full description on the dataset page: https://huggingface.co/datasets/ankitha29/telugu-colloquial-corpus-tokenized.telugu-colloquial-corpus-tokenized
Telugu Colloquial Corpus (Tokenized)
This dataset is a tokenized version of the Telugu Colloquial Corpus (TeCC). It contains examples of informal, everyday Telugu language, including slang, regional variations, and conversational patterns.
Dataset Details
Language: Telugu (te)
Tokenization: Tokenized using the bert-base-multilingual-cased tokenizer from the transformers library.
Source: [Describe where the original data came from – e.g., collected from online forums… See the full description on the dataset page: https://huggingface.co/datasets/dokkuayyappa5/telugu-colloquial-corpus-tokenized.titulm-bangla-corpus-tokenizedvngrs-web-corpus-200k-kumru_tokenizer-tokenizedminicpm5-agent-corpus-8k-tokenizedgpt2_small_pretraining_corpus_tokenized_256sozkz-corpus-tokenized-kk-morph-v1sozkz-corpus-tokenized-enkk-fineweb-edu-v2sozkz-corpus-tokenized-kk-t5-50m-v1sozkz-corpus-tokenized-kk-morphbpe100k-v1
