CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jinofy-corp /jora_corpus1_tokenized_128ktabularn<1K4 likes42k downloads2mo agoHugging Face02AINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline. Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes865 downloads23d agoHugging Face03ynnekuw /common-corpus-noncc-tokenized10M<n<100M0 likes793 downloads3mo agoHugging Face04jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes614 downloads2mo agoHugging Face05nicholasKluge /Pt-Corpus-Instruct-tokenized Portuguese-Corpus Instruct (tokenized) Dataset Summary This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese". For more information, see the original dataset card. Languages Portuguese. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.text-generation1M<n<10M0 likes584 downloads1y agoHugging Face06Ba2han /tokenized-corpus-0603 Tokenized Corpus Statistics Shard Total Tokens Avg Tokens Median Tokens train-00000-of-00077 64,973,050 259.89 213.0 train-00001-of-00077 65,024,877 260.10 214.0 train-00002-of-00077 64,976,757 259.91 213.0 train-00003-of-00077 65,035,795 260.14 214.0 train-00004-of-00077 64,962,798 259.85 214.0 train-00005-of-00077 64,945,077 259.78 213.0 train-00006-of-00077 64,885,161 259.54 213.0 train-00007-of-00077 65,060,639 260.24 214.0 train-00008-of-0007764,905… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/tokenized-corpus-0603.10M<n<100M0 likes484 downloads7mo agoHugging Face07umarigan /turkish_corpus_tokenized Dataset Card for "turkish_corpus_tokenized" More Information needed 10M<n<100M0 likes355 downloads3y agoHugging Face08Kashif786 /sindhi-gold-corpus-mlm-tokenized-for-bert1M<n<10M0 likes333 downloads27d agoHugging Face09philschmid /llama2-german-corpus-tokenized-llama-chunk-4096 Dataset Card for "llama2-german-corpus-tokenized-llama-chunk-4096" More Information needed 10M<n<100M0 likes300 downloads3y agoHugging Face10nicholasKluge /Pt-Corpus-tokenized Portuguese-Corpus (tokenized) Dataset Summary This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus dataset. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese". For more information, see the original dataset card. Languages Portuguese. Dataset Structure Data Instances The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-tokenized.text-generation1M<n<10M0 likes224 downloads2y agoHugging Face11Ba2han /long_corpus-0209_tokenized long_corpus-0209_tokenized Exact-deduplicated, tokenized, and shuffled pretraining mix. Processing Tokenizer: /workspace/good_tokenizer (vocab_size=60800) Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3) bos_eos = True: every example is [BOS] + content + [EOS] min_tokens = 15 (including BOS/EOS) max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos) Dedup: exact match on stripped UTF-8 text (blake2s-128) Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.text-generation10M<n<100M0 likes172 downloads19d agoHugging Face12ZurabDz /tokenized_large_corpus_v2 Dataset Card for "tokenized_large_corpus_v2" More Information needed 10M<n<100M0 likes164 downloads3y agoHugging Face13trungbb8 /vietnamese-news-corpus-tokenized10M<n<100M0 likes151 downloads4mo agoHugging Face14Kashif786 /sindhi-gold-corpus-mlm-tokenized-mbert1M<n<10M0 likes79 downloads18d agoHugging Face15chenrm /h-corpus-tokenized-qwen2.5-ctx20481M<n<10M0 likes75 downloads2y agoHugging Face16mrm8488 /large_spanish_corpus_ds_tokenized_and_gropuped Dataset Card for "large_spanish_corpus_ds_tokenized_and_gropuped" More Information needed 1M<n<10M0 likes59 downloads4y agoHugging Face17Veenn /magnetar-tokenized-corpusgated0 likes56 downloads15d agoHugging Face18Abzalbek89 /corpus_clean_tokenized Kazakh Tokenized Corpus (2048 blocks) Pre-tokenized Kazakh corpus ready for language model training. Built from Abzalbek89/corpus_clean using Abzalbek89/kk-tokenizer-bpe-32k. Dataset Summary Metric Value Train blocks 236,981 Validation blocks 12,473 Total blocks 249,454 Block size 2,048 tokens Total tokens ~0.51B (510M) Tokenizer ByteLevel BPE, 32K vocab Val ratio 5% Pipeline Source: Abzalbek89/corpus_clean — 1,502,583 cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Abzalbek89/corpus_clean_tokenized.text-generation100K<n<1M0 likes52 downloads5mo agoHugging Face19ynnekuw /common-corpus-noncc-500b-tokenized10M<n<100M0 likes48 downloads2mo agoHugging Face20canbingol /vngrs-web-corpus-500k-kumru_tokenizer-tokenized0 likes12 downloads8mo agoHugging Face21ankitha29 /telugu-colloquial-corpus-tokenized Telugu Colloquial Corpus (Tokenized) This dataset is a tokenized version of the Telugu Colloquial Corpus (TeCC). It contains examples of informal, everyday Telugu language, including slang, regional variations, and conversational patterns. Dataset Details Language: Telugu (te) Tokenization: Tokenized using the bert-base-multilingual-cased tokenizer from the transformers library. Source: [Describe where the original data came from – e.g., collected from online forums… See the full description on the dataset page: https://huggingface.co/datasets/ankitha29/telugu-colloquial-corpus-tokenized.text-generationn<1K0 likes9 downloads2y agoHugging Face22dokkuayyappa5 /telugu-colloquial-corpus-tokenized Telugu Colloquial Corpus (Tokenized) This dataset is a tokenized version of the Telugu Colloquial Corpus (TeCC). It contains examples of informal, everyday Telugu language, including slang, regional variations, and conversational patterns. Dataset Details Language: Telugu (te) Tokenization: Tokenized using the bert-base-multilingual-cased tokenizer from the transformers library. Source: [Describe where the original data came from – e.g., collected from online forums… See the full description on the dataset page: https://huggingface.co/datasets/dokkuayyappa5/telugu-colloquial-corpus-tokenized.text-generationn<1K0 likes9 downloads7mo agoHugging Face23spitfire4794 /titulm-bangla-corpus-tokenizedtabularn<1K0 likes9 downloads3mo agoHugging Face24canbingol /vngrs-web-corpus-200k-kumru_tokenizer-tokenized0 likes8 downloads8mo agoHugging Face25dschauhan08 /minicpm5-agent-corpus-8k-tokenizedtabular100K<n<1M0 likes7 downloads2mo agoHugging Face26pritamdeb68 /gpt2_small_pretraining_corpus_tokenized_256100K<n<1M0 likes6 downloads8mo agoHugging Face27stukenov /sozkz-corpus-tokenized-kk-morph-v1gated100K<n<1M0 likes6 downloads7mo agoHugging Face28stukenov /sozkz-corpus-tokenized-enkk-fineweb-edu-v2gated1M<n<10M0 likes4 downloads6mo agoHugging Face29stukenov /sozkz-corpus-tokenized-kk-t5-50m-v1gated1M<n<10M0 likes3 downloads7mo agoHugging Face30stukenov /sozkz-corpus-tokenized-kk-morphbpe100k-v1gated1M<n<10M0 likes3 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.