CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anisoleai /fineweb-tokenized FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.tabulartext-generationn>1T34 likes42k downloads4mo agoHugging Face02TrevorDohm /Stack_Tokenizedtexttext-generation100M<n<1B0 likes3k downloads2y agoHugging Face03mondk /fineweb-tokenized-fake What is it? It's similar to anisolai/fineweb-tokenized but fake. I don't understand why I did that :) WARNING: WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.tabulartext-generation10M<n<100M2 likes1.3k downloads25d agoHugging Face04AethronPhantom /Scientific_Research_Tokenized NexaSci Scientific Research Tokenized This dataset repository now holds the active NexaSci scientific pretraining reservoir, the NexaMat controller fine-tuning pack, and archived legacy reservoir builds. The current production reservoir is the 10B-token Apache Arrow release under nexasci_reservoir_v3_10b_prod_rust/. Current Status The active large-scale training artifact is: nexasci_reservoir_v3_10b_prod_rust/ It was produced from the NexaSci 10B data-engineering campaign… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Scientific_Research_Tokenized.texttext-generation100K<n<1M7 likes1.2k downloads4mo agoHugging Face05LLM-OS-Models /Qwen-Terminal-ToolBench-Processed-Tokenized Qwen Terminal ToolBench Processed Datasets Qwen-family processed/template-applied and selected tokenized terminal datasets. Contents qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.text-generation0 likes1k downloads4mo agoHugging Face06placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes883 downloads1mo agoHugging Face07AINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized representation of sequence intervals processed in the Carbon enrichment pipeline. Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes865 downloads24d agoHugging Face08abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes726 downloads2mo agoHugging Face09Sambarboi /climbmix-tokenized-20480-diloco ClimbMix, retokenized and shuffled for three-worker DiLoCo This is a document-preserving, three-way split of NVIDIA's Nemotron-ClimbMix, retokenized with a 20,480-entry byte-level BPE tokenizer. Each document ends in <|endoftext|>. The Arrow IPC streams use transparent Zstandard buffer compression. A deterministic whole-shard holdout is shared by every worker for validation and is excluded from training. Training part Documents Tokens Files Compressed size 000 15,709… See the full description on the dataset page: https://huggingface.co/datasets/Sambarboi/climbmix-tokenized-20480-diloco.text-generation1M<n<10M0 likes623 downloads2mo agoHugging Face10ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes621 downloads17d agoHugging Face11nicholasKluge /Pt-Corpus-Instruct-tokenized Portuguese-Corpus Instruct (tokenized) Dataset Summary This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus Instruct dataset. All sequences are 2048 tokens long. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese". For more information, see the original dataset card. Languages Portuguese. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct-tokenized.text-generation1M<n<10M0 likes584 downloads1y agoHugging Face12emozilla /dolma-v1_7-305B-tokenized-llama3-nanosetTokenized (Llama 3) verison of NousResearch/dolma-v1_7-305B as a Nanotron dataset split into 10 GB chunks. To download: huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-305B-tokenized-llama3-nanoset --local-dir-use-symlinks False NousResearch/dolma-v1_7-305B-tokenized-llama3-nanoset To recombine: cat dolma-v1_7-305B-tokenized-llama3-nanoset/dolma-v1_7-305B-tokenized-llama3-nanoset.npy.* > dolma-v1_7-305B-tokenized-llama3-nanoset.npy rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-305B-tokenized-llama3-nanoset.text-generation100B<n<1T1 likes461 downloads2y agoHugging Face13placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes418 downloads1mo agoHugging Face14zee-drytis /damr-zyda2-64k-tokenized DAMR Zyda2 64K Tokenized This repository contains a deterministic tokenized representation of a 1,749,978,785,641-token subset of Zyda-2 for language-model pretraining. Format Files under data/ contain contiguous little-endian uint16 token IDs. Concatenate shards in numeric order to reproduce the original stream. The shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard manifests provide byte offsets, token counts, and SHA-256 checksums. The… See the full description on the dataset page: https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.text-generation1 likes414 downloads2mo agoHugging Face15nikolina-p /fineweb_10BT_tokenized Dataset card for FineWeb-Edu 10B tokenized dataset This dataset contains tokenized texts from FineWeb-Edu sample-10B HuggingFaceFW/fineweb-edu. The data was tokenized using the OpenAI's tiktoken tokenizer, and structured for efficient streaming and distributed (DDP) training. Structure The dataset follows Hugging Face’s recommended structure for efficient streaming in multi-GPU environments. It consists of two splits, where each split contains a number of shards… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/fineweb_10BT_tokenized.text-generation10K<n<100K0 likes330 downloads11mo agoHugging Face16izlley2 /llm0to1-pt-tokenized-en-edu-2025 LLM0to1 사전학습 토큰화본 — 영어 교육(2025 덤프) 10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된 영어 교육(2025 덤프) 코퍼스의 토큰화본. 총 28종 / 138.4B 토큰. 왜 원문 텍스트가 아니라 토큰화본인가 이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다. 따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다. 단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로, 재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다. 원본 출처 HuggingFaceFW/fineweb-edu 2025 덤프 영어 벤치마크 하락에 대응해 뒤늦게 편입한 '다양성 영어' 보강분이다. g00~`g15` 는 원본 shard 를 균등 분할한 것으로, 서로 다른 문서 집합이다.… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-en-edu-2025.text-generationn>1T0 likes329 downloads1mo agoHugging Face17CausalNLP /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes320 downloads2mo agoHugging Face18izlley2 /llm0to1-pt-tokenized-code LLM0to1 사전학습 토큰화본 — 코드 10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된 코드 코퍼스의 토큰화본. 총 27종 / 59.1B 토큰. 왜 원문 텍스트가 아니라 토큰화본인가 이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다. 따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다. 단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로, 재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다. 원본 출처 bigcode/starcoderdata(언어별 서브셋) · bigcode/commitpackft · bigcode/jupyter-code-text-pairs · deepmind/code_contests code_c 와 code_c2 처럼 2 가 붙은 것은 정제기… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-code.text-generationn>1T0 likes298 downloads1mo agoHugging Face19placeholderlabs /exp-pool-academic-dolma2-tokenized Locus EXP Academic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes295 downloads1mo agoHugging Face20birgermoell /oellm-longctx-tokenized-streamed-all-v2 OpenEuroLLM long-context Megatron streamed tokenized artifacts This dataset contains Megatron-LM indexed data (.bin/.idx) uploaded shard by shard from longctx stream-upload. It is intended as a transport format for continued pretraining on LUMI or another training machine; it is not raw text. Source dataset: HuggingFaceFW/finepdfs-edu Tooling: openeuro-longctx-datamix Run namespace: runs/lc16k_full_20260507/ Tokenizer type: HuggingFaceTokenizer Tokenizer: OpenEuroLLM 256k… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-streamed-all-v2.text-generation0 likes278 downloads4mo agoHugging Face21SauravP97 /tiny-stories-tokenized-bpetexttext-generation1M<n<10M1 likes246 downloads7mo agoHugging Face22placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes243 downloads1mo agoHugging Face23YangyiH /DepthBench-FineWeb-Edu-100BT-tokenized DepthBench FineWeb-Edu 100BT Tokenized This repository contains the tokenized FineWeb-Edu 100BT sample used by DepthBench pretraining experiments. Splits train/: 139 shards, 99,585,913,529 tokens, and 97,045,608 documents. eval/: 013_00008, containing 234,993,701 tokens and 225,078 documents. All remaining source shards are assigned to training. Each source document is terminated by an EOS token before documents are concatenated. Format Each shard… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/DepthBench-FineWeb-Edu-100BT-tokenized.tabulartext-generationn<1K0 likes225 downloads2mo agoHugging Face24nicholasKluge /Pt-Corpus-tokenized Portuguese-Corpus (tokenized) Dataset Summary This repository has a tokenized version (using the TeenyTinyLlama tokenizer) of the Portuguese-Corpus dataset. All sequences are 2048 tokens long. This dataset was used in "TeenyTinyLlama: open-source tiny language models trained in Brazilian Portuguese". For more information, see the original dataset card. Languages Portuguese. Dataset Structure Data Instances The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-tokenized.text-generation1M<n<10M0 likes224 downloads2y agoHugging Face25placeholderlabs /exp-pool-encyclopedic-dolma2-tokenized Locus EXP Encyclopedic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.tabulartext-generation1M<n<10M0 likes220 downloads1mo agoHugging Face26placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes210 downloads1mo agoHugging Face27Ba2han /long_corpus-0209_tokenized long_corpus-0209_tokenized Exact-deduplicated, tokenized, and shuffled pretraining mix. Processing Tokenizer: /workspace/good_tokenizer (vocab_size=60800) Special tokens: BOS=<bos> (id 2), EOS=<eos> (id 3) bos_eos = True: every example is [BOS] + content + [EOS] min_tokens = 15 (including BOS/EOS) max_tokens = 3072: content truncated to 3070 tokens, then BOS+EOS appended (3070 + eos + bos) Dedup: exact match on stripped UTF-8 text (blake2s-128) Shuffle:… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/long_corpus-0209_tokenized.text-generation10M<n<100M0 likes172 downloads20d agoHugging Face28andreuka18 /DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenizedThis dataset is used for training Sparse Autoencoders (SAEs) to identify reasoning features in Large Language Models (LLMs), as described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders. Code for the paper is available at: https://github.com/AIRI-Institute/SAE-Reasoning The dataset consists of tokenized text data used for training the SAEs. dataset_info: features: name: tokens sequence: int64 splits: name:… See the full description on the dataset page: https://huggingface.co/datasets/andreuka18/DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized.text-generation100K<n<1M0 likes160 downloads1y agoHugging Face29CodeForCodersYT /ultimate-code-tokenized ultimate-code Dataset Description ultimate-code is a derived dataset built by combining and processing data from nvidia/OpenCodeInstruct and nvidia/OpenCodeGeneticInstruct. It can be used to fine-tune LLMs for coding tasks. Tokenized variant available: a pre-tokenized version of this dataset is available at CodeForCodersYT/ultimate-code-tokenized. Use that version if you want ready-to-train tokenized sequences instead of raw text. Source Datasets &… See the full description on the dataset page: https://huggingface.co/datasets/CodeForCodersYT/ultimate-code-tokenized.texttext-generation10M<n<100M0 likes155 downloads2mo agoHugging Face30Neetree /fineweb10B-tokenized-custom LittleTzu FineWeb-Edu Tokenized (Custom 65k Balanced) Tokenized shards of FineWeb-Edu (HuggingFaceFW/fineweb-edu, config: sample-10BT) for language model pretraining. This dataset stores a derived, tokenized representation of the original FineWeb-Edu corpus. It has been tokenized using LittleTzu's custom 65K balanced tokenizer, optimized for multi-domain training (English, multilingual text, math, and code) while maintaining a compact vocabulary footprint that fits within a… See the full description on the dataset page: https://huggingface.co/datasets/Neetree/fineweb10B-tokenized-custom.text-generation10B<n<100B0 likes152 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.