CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anisoleai /fineweb-tokenized FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.tabulartext-generationn>1T34 likes38k downloads4mo agoHugging Face02asahi417 /seamless-align-enA-jaA.tokenized.encodectabular100K<n<1M0 likes1.7k downloads2y agoHugging Face03mondk /fineweb-tokenized-fake What is it? It's similar to anisolai/fineweb-tokenized but fake. I don't understand why I did that :) WARNING: WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.tabulartext-generation10M<n<100M2 likes1.5k downloads27d agoHugging Face04asahi417 /seamless-align-enA-hiA.tokenized.encodectabular100K<n<1M0 likes1.3k downloads2y agoHugging Face05asahi417 /seamless-align-enA-esA.tokenized.encodectabular100K<n<1M0 likes1.1k downloads2y agoHugging Face06asahi417 /seamless-align-enA-viA.tokenized.encodectabular100K<n<1M0 likes1.1k downloads2y agoHugging Face07asahi417 /seamless-align-deA-enA.tokenized.encodectabular100K<n<1M0 likes1k downloads2y agoHugging Face08AINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized sample of AINovice2005/carbon-cpu-enriched-sequences-sampled . Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer token_mask list… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes875 downloads2d agoHugging Face09argo11 /0399-tv-valid-clean-sft-tokenized-llmjp4-8btabular1M<n<10M0 likes843 downloads3mo agoHugging Face10placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes664 downloads1mo agoHugging Face11Pclanglais /tokenized_sampletabular1M<n<10M0 likes620 downloads2y agoHugging Face12asahi417 /seamless-align-enA-koA.tokenized.encodectabular100K<n<1M0 likes563 downloads2y agoHugging Face13asahi417 /seamless-align-enA-frA.tokenized.encodectabular1M<n<10M0 likes563 downloads2y agoHugging Face14Menlo /Ichigo-instruction-tokenized-v0.2tabular1M<n<10M0 likes548 downloads2y agoHugging Face15mikaberidze /sib200-xlmr-tokenized SIB-200 Tokenized by XLM-R Large This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352. The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper, ensuring consistent and reproducible tokenization across all languages and experimental settings. The dataset is organized as a multi-config Hugging… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.tabulartext-classification100K<n<1M0 likes490 downloads9mo agoHugging Face16aklein4 /mixed-pretraining-tokenizedtabular10M<n<100M0 likes427 downloads1y agoHugging Face17mjbommar /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.tabularother10K<n<100K1 likes414 downloads11mo agoHugging Face18KrisMinchev /finemath-4plus-tokenizedtabular1M<n<10M0 likes411 downloads9mo agoHugging Face19deatos /tokenized_fineweb_edu_10b_combinedtabular1M<n<10M0 likes394 downloads2y agoHugging Face20asahi417 /seamless-align-enA-zhA.tokenized.encodectabular100K<n<1M0 likes378 downloads2y agoHugging Face21TemryL /tokenized_wikipedia_20220301.en_train_512 Tokenized English Wikipedia Dataset Dataset Description This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from. Dataset Creation Source Dataset: Wikipedia (20220301.en) Tokenizer: BERT base uncased Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.tabular10M<n<100M0 likes319 downloads2y agoHugging Face22aklein4 /HuggingFaceFW-fineweb-sample-10BT-tokenizedtabular10M<n<100M0 likes305 downloads1y agoHugging Face23Menlo /Ichigo-instruction-tokenized-v0.2-cleantabular1M<n<10M0 likes285 downloads2y agoHugging Face24SemantikaEU /MaCoCu-sl-tokenized Dataset Card for MaCoCu-sl Multi-Tokenized Dataset Description: This dataset provides a pre-tokenized version of the Slovene web corpus MaCoCu. It includes the original text data and metadata from MaCoCu-sl, augmented with token IDs and token counts generated by several popular large language model tokenizers. The goal is to facilitate research and experimentation by providing ready-to-use tokenized data, saving computational resources during repeated setups. Licensing and… See the full description on the dataset page: https://huggingface.co/datasets/SemantikaEU/MaCoCu-sl-tokenized.tabular1M<n<10M0 likes277 downloads1y agoHugging Face25placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes239 downloads1mo agoHugging Face26aklein4 /HuggingFaceFW-fineweb-edu-sample-10BT-tokenizedtabular1M<n<10M0 likes225 downloads1y agoHugging Face27FundamentalResearchLabs /leader-training-tokenized-fixed-summary-16k-fulltabular100K<n<1M0 likes201 downloads1y agoHugging Face28aklein4 /HuggingFaceTB-finemath-finemath-4plus-tokenizedtabular1M<n<10M0 likes196 downloads1y agoHugging Face29placeholderlabs /exp-pool-academic-dolma2-tokenized Locus EXP Academic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes176 downloads1mo agoHugging Face30Menlo /Ichigo-instruction-tokenized-v0.1tabular1M<n<10M0 likes160 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.