CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01anisoleai /fineweb-tokenized FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.tabulartext-generationn>1T34 likes37k downloads4mo agoHugging Face02open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes2.1k downloads2y agoHugging Face03asahi417 /seamless-align-enA-jaA.tokenized.encodectabular100K<n<1M0 likes1.7k downloads2y agoHugging Face04mondk /fineweb-tokenized-fake What is it? It's similar to anisolai/fineweb-tokenized but fake. I don't understand why I did that :) WARNING: WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE. WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.tabulartext-generation10M<n<100M2 likes1.6k downloads28d agoHugging Face05asahi417 /seamless-align-enA-hiA.tokenized.encodectabular100K<n<1M0 likes1.3k downloads2y agoHugging Face06ceselder /loracle-pretrain-v5-qwen14b-tokenstabular10K<n<100K0 likes1.2k downloads2mo agoHugging Face07SaylorTwift /RULER-8192-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes1.1k downloads1y agoHugging Face08asahi417 /seamless-align-enA-esA.tokenized.encodectabular100K<n<1M0 likes1.1k downloads2y agoHugging Face09asahi417 /seamless-align-enA-viA.tokenized.encodectabular100K<n<1M0 likes1.1k downloads2y agoHugging Face10asahi417 /seamless-align-deA-enA.tokenized.encodectabular100K<n<1M0 likes1k downloads2y agoHugging Face11AINovice2005 /carbon-tokenized-corpus Dataset Summary AINovice2005/carbon-tokenized-corpus is the tokenized sample of AINovice2005/carbon-cpu-enriched-sequences-sampled . Schema The current dataset contains the following fields: Field Type Description record_id string Source/reference sequence identifier start int64 Start coordinate of the sequence interval end int64 End coordinate of the sequence interval token_ids list Integer token IDs produced by the tokenizer token_mask list… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.tabulartext-generation1M<n<10M0 likes844 downloads2d agoHugging Face12argo11 /0399-tv-valid-clean-sft-tokenized-llmjp4-8btabular1M<n<10M0 likes826 downloads3mo agoHugging Face13BlockDB /Token-To-Token-Prices-OHLC-Ethereum-Cryptocurrency-Data Token-To-Token-Prices-OHLC-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB token_to_token_prices_ohlc (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/Token-To-Token-Prices-OHLC-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/Token-To-Token-Prices-OHLC-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes794 downloads29d agoHugging Face14BlockDB /Token-To-Token-VWAP-Ethereum-Cryptocurrency-Data Token-To-Token-VWAP-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB token_to_token_vwap (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/Token-To-Token-VWAP-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp bucket_start timestamp… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/Token-To-Token-VWAP-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes731 downloads29d agoHugging Face15placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes662 downloads1mo agoHugging Face16Pclanglais /tokenized_sampletabular1M<n<10M0 likes611 downloads2y agoHugging Face17BlockDB /Token-To-Token-Cross-Pool-VWAP-Ethereum-Cryptocurrency-Data Token-To-Token-Cross-Pool-VWAP-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB token_to_token_cross_pool_vwap (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/Token-To-Token-Cross-Pool-VWAP-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/Token-To-Token-Cross-Pool-VWAP-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes607 downloads29d agoHugging Face18Menlo /Ichigo-instruction-tokenized-v0.2tabular1M<n<10M0 likes547 downloads2y agoHugging Face19asahi417 /seamless-align-enA-frA.tokenized.encodectabular1M<n<10M0 likes545 downloads2y agoHugging Face20asahi417 /seamless-align-enA-koA.tokenized.encodectabular100K<n<1M0 likes537 downloads2y agoHugging Face21BlockDB /ERC20-Tokens-Ethereum-Cryptocurrency-Data ERC20-Tokens-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB erc20_tokens (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/ERC20-Tokens-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2023-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp block_number int64 tx_index int32 contract_id… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/ERC20-Tokens-Ethereum-Cryptocurrency-Data.tabular1M<n<10M0 likes523 downloads29d agoHugging Face22BlockDB /Token-To-Token-Prices-Swap-Prints-Ethereum-Cryptocurrency-Data Token-To-Token-Prices-Swap-Prints-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB token_to_token_prices_swap_prints (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/Token-To-Token-Prices-Swap-Prints-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2020-05 .. 2026-06 (UTC calendar months). Schema column type block_timestamp… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/Token-To-Token-Prices-Swap-Prints-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes506 downloads1mo agoHugging Face23valentinhofmann /c4-token-logprobstabular1M<n<10M0 likes505 downloads21d agoHugging Face24SaylorTwift /RULER-32768-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes491 downloads1y agoHugging Face25mikaberidze /sib200-xlmr-tokenized SIB-200 Tokenized by XLM-R Large This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352. The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper, ensuring consistent and reproducible tokenization across all languages and experimental settings. The dataset is organized as a multi-config Hugging… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.tabulartext-classification100K<n<1M0 likes489 downloads9mo agoHugging Face26BlockDB /Token-To-Fiat-VWAP-Ethereum-Cryptocurrency-Data Token-To-Fiat-VWAP-Ethereum-Cryptocurrency-Data Hive-partitioned Parquet export of BlockDB token_to_fiat_vwap (Ethereum). Load from datasets import load_dataset ds = load_dataset("BlockDB/Token-To-Fiat-VWAP-Ethereum-Cryptocurrency-Data", split="train") Files live under data/year=YYYY/month=MM/part-NNNN.parquet. Range: 2015-08 .. 2026-06 (UTC calendar months). Schema column type block_timestamp timestamp bucket_start timestamp… See the full description on the dataset page: https://huggingface.co/datasets/BlockDB/Token-To-Fiat-VWAP-Ethereum-Cryptocurrency-Data.tabular100M<n<1B0 likes445 downloads29d agoHugging Face27kothasuhas /dclm_10B_tokenstabular1M<n<10M0 likes436 downloads1y agoHugging Face28aklein4 /mixed-pretraining-tokenizedtabular10M<n<100M0 likes433 downloads1y agoHugging Face29KrisMinchev /finemath-4plus-tokenizedtabular1M<n<10M0 likes413 downloads9mo agoHugging Face30asahi417 /seamless-align-enA-zhA.tokenized.encodectabular100K<n<1M0 likes394 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.