CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01occiglot /tokenizer-wiki-bench Multilingual Tokenizer Benchmark This dataset includes pre-processed wikipedia data for tokenizer evaluation in 45 languages. We provide more information on the evaluation task in general this blogpost. Usage The dataset allows us to easily calculate tokenizer fertility and the proportion of continued words on any of the supported languages. In the example below we take the Mistral tokenizer and evaluate its performance on Slovak. from transformers import AutoTokenizer… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/tokenizer-wiki-bench.text10M<n<100M6 likes48k downloads2y agoHugging Face02jinofy-corp /jora_corpus1_tokenized_128ktabularn<1K4 likes40k downloads2mo agoHugging Face03anisoleai /fineweb-tokenized FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.tabulartext-generationn>1T34 likes37k downloads4mo agoHugging Face04hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes27k downloads14d agoHugging Face05tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes13k downloads11mo agoHugging Face06AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes9.4k downloads3y agoHugging Face07tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B48 likes8.3k downloads11mo agoHugging Face08syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face09TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes5.6k downloads3y agoHugging Face10Biomedical-TeMU /SPACCC_Tokenizer The Tokenizer for Clinical Cases Written in Spanish Introduction This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish. This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.text10K<n<100K0 likes4.5k downloads5y agoHugging Face11napaull /tokenized_C4textn<1K0 likes3.3k downloads5mo agoHugging Face12TrevorDohm /Pile_TokLlama Dataset Card for "Pile_TokLlama" More Information needed text100M<n<1B0 likes3k downloads2y agoHugging Face13ppbrown /tokenspace tokenspace directory This directory contains utilities for the purpose of browsing the "token space" of CLIP ViT-L/14 Primary tools are: "calculate-distances.py": allows command-line browsing of words and their neighbours "graph-embeddings.py": plots graph of full values of two embeddings (clipmodel,cliptextmodel)-calculate-distances.py Loads the generated embeddings, reads in a word, calculates "distance" to every embedding, and then shows the closest "neighbours". To… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/tokenspace.textn<1K7 likes3k downloads2y agoHugging Face14TrevorDohm /Stack_Tokenizedtexttext-generation100M<n<1B0 likes2.8k downloads2y agoHugging Face15TokenRhythm /Claw-SWE-Bench Claw-SWE-Bench Paper: Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-Style Agent Harnesses on Coding Tasks A multilingual issue-resolving benchmark with two evaluation configs: full — 350 instances (300 from SWE-bench Multilingual + 50 Python from SWEBench-verified-mini's size_optimized_sample). lite — 80-instance calibrated subset (10 per language across 8 languages: Java, Go, Rust, JS/TS, C/C++, Ruby, PHP, Python). Designed for low-cost iteration on harness… See the full description on the dataset page: https://huggingface.co/datasets/TokenRhythm/Claw-SWE-Bench.texttext-generationn<1K8 likes2.8k downloads4mo agoHugging Face16upup-ashton-wang /temp-bert-train-tokenizedtextn<1K0 likes2.3k downloads5mo agoHugging Face17open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face18daje /tokenized_enwiki Dataset Card for "tokenized_enwiki" More Information needed text10M<n<100M0 likes2k downloads3y agoHugging Face19arolstar52 /ocr-synthetic-multilingual-v1-tokenized-zh-hanstext1M<n<10M0 likes1.9k downloads17d agoHugging Face20weikaih /imaginative-perception-token-pet-ipt Citation Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988): @misc{bigverdi2026imaginativeperceptiontokensenhance, title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models}, author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pet-ipt.image10K<n<100K1 likes1.9k downloads4mo agoHugging Face21krvhrv /Healix-2.8B-Token-Medical-Shot Dataset Card for "Healix-2.8B-Token-Medical-Shot" More Information needed text1M<n<10M0 likes1.7k downloads3y agoHugging Face22asahi417 /seamless-align-enA-jaA.tokenized.encodectabular100K<n<1M0 likes1.7k downloads2y agoHugging Face23Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.7k downloads5mo agoHugging Face24andersonbcdefg /PD-3M-Tokenized-Cosmos-Tokenizer-DI8x8I can't get the dataset viewer to work, sorry. There's about 3M images and captions from Spawning/PD3M. They are resized and center-cropped to 512x512, and then tokenized into discrete tokens with NVIDIA Cosmos-Tokenizer-DI8x8, which reduces the spatial dimension by a factor of 8, resulting in 64 x 64 = 4096 discrete tokens per image. You can use these tokenized images to train an auto-regressive image model, or a MaskGIT. Or probably other things I don't know about. :) License is the same… See the full description on the dataset page: https://huggingface.co/datasets/andersonbcdefg/PD-3M-Tokenized-Cosmos-Tokenizer-DI8x8.text10M<n<100M0 likes1.6k downloads2y agoHugging Face25brendanlong /subliminal-transfer-token-replacement Subliminal transfer: token replacement vs masking (artifacts) Teachers, training data, per-token divergence scores and evaluation outputs for brendanlong/subliminal-transfer-token-replacement. The experiment asks whether replacing attribution-flagged tokens suppresses a subliminally transmitted trait better than masking them from the loss, and whether any advantage is specific to those tokens. Everything here is for the one studied cell: Llama-3.2-1B-Instruct, target animal… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/subliminal-transfer-token-replacement.tabulartext-generationn<1K0 likes1.6k downloads6d agoHugging Face26catherinearnett /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B1 likes1.6k downloads1y agoHugging Face27mondk /fineweb-tokenized-fake What is it? It's similar to anisolai/fineweb-tokenized but fake. I don't understand why I did that :) WARNING: WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.tabulartext-generation10M<n<100M2 likes1.6k downloads28d agoHugging Face28weikaih /imaginative-perception-token-mvc-ipt Citation Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988): @misc{bigverdi2026imaginativeperceptiontokensenhance, title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models}, author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-mvc-ipt.image10K<n<100K1 likes1.5k downloads4mo agoHugging Face29chainyo /natural-instructions-tokenized Dataset Card for "natural-instructions-tokenized" Here is the script used to tokenize the dataset: import multiprocessing from typing import Union from datasets import DatasetDict, load_dataset from transformers import LlamaTokenizer # Find your available cores num_cores = multiprocessing.cpu_count() cutoff_len = 2048 tokenizer = LlamaTokenizer.from_pretrained("chainyo/alpaca-lora-7b") tokenizer.padding_side = "left" tokenizer.pad_token_id = (0) prompt_template = {… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/natural-instructions-tokenized.text1M<n<10M1 likes1.5k downloads3y agoHugging Face30arolstar52 /ocr-synthetic-multilingual-v1-tokenized-entext1M<n<10M0 likes1.5k downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.