CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ks46 /urls-tokenized URLs (tokenized) ks46/urls-sampled run through a byte-level BPE built for URLs, stored as flat uint16 token streams that memory-map directly into a training loop. Shards 512 URLs 18,729,786,698 Tokens 664,731,047,208 Vocabulary 8,192 Token dtype uint16, little-endian There is no parquet here and the dataset viewer will not render it. These are raw token bins; see Reading the data below. Layout tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tabulartext-generationn<1K0 likes622 downloads21d agoHugging Face02erickfmm /tokenised_subsetof_erickfmm__red_pajama_es_hq_35tabularn<1K0 likes592 downloads8mo agoHugging Face03Efe2898 /tokenizedtabularn<1K0 likes383 downloads3d agoHugging Face04eduagarcia /multilingual_tokenizer_benchmark Multilingual Tokenizer Benchmark More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root. Natural language word count functions Download spacy models pip install ntlk spacy pygments underthesea camel-tools python -m spacy download ko_core_news_sm python -m spacy download ja_core_news_sm python -m spacy download zh_core_web_sm import nltk nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.tabulartext-generation100K<n<1M2 likes224 downloads1y agoHugging Face05Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes196 downloads26d agoHugging Face06jinofy-corp /jora_corpus1_FR_tokenized_128ktabularn<1K2 likes188 downloads2mo agoHugging Face07procmarco /fpabl1-arm-b-fp-tokens-48k fpabl1-arm-b-fp-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fp_en: 1,000,000,000 fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.tabularn<1K0 likes184 downloads3mo agoHugging Face08zjhhhh /fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes140 downloads21d agoHugging Face09soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-03-token-cachetabularn<1K0 likes101 downloads1mo agoHugging Face10Pacific-i64 /data-32k-200b-tokens TR-HASH 32K · 200B Token Mixture Pretokenized training mixture for compact TR-HASH language-model research. The Dataset Viewer displays one summary row per source. The actual training data is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin. Each sequence contains 1,024 token IDs produced by the project 32K tokenizer. Mixture Source Weight Training tokens DCLM 45% 90B FineWeb-Edu deduplicated 30% 60B Stack-Edu 10% 20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.tabulartext-generationn<1K1 likes99 downloads1mo agoHugging Face11soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-02-token-cachetabularn<1K0 likes89 downloads1mo agoHugging Face12zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes85 downloads20d agoHugging Face13soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-01-token-cachetabularn<1K0 likes78 downloads1mo agoHugging Face14soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-layer-00-token-cachetabularn<1K0 likes72 downloads1mo agoHugging Face15zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes72 downloads1mo agoHugging Face16soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s1-token-cachetabularn<1K0 likes71 downloads1mo agoHugging Face17soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2308-s2-token-cachetabularn<1K0 likes67 downloads1mo agoHugging Face18kenny0bi /tokenizer-tax owóorí: the tokenizer tax, measured Token-count premiums for the 204 languages of FLORES-200 under 12 tokenizers, from the identical 1,012 professionally translated sentences. The premium is tokens(language) / tokens(English) on the same content, so it reads directly as a price multiplier for API cost, latency and context shrinkage. Interactive explorer: https://kenny0bi.github.io/owoori/ Method, figures, code: https://github.com/Kenny0bi/owoori Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.tabularn<1K0 likes66 downloads24d agoHugging Face19soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s2-token-cachetabularn<1K0 likes63 downloads1mo agoHugging Face20toksuitebackup /tokenmonster-englishcode-32000-consistent-v1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes60 downloads10mo agoHugging Face21soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-2400-s0-token-cachetabularn<1K0 likes60 downloads1mo agoHugging Face22soar-eleuther-i6-hierarchy /metrics-outputs-gemma-2-2b-layer-06-token-cachetabularn<1K0 likes60 downloads25d agoHugging Face23birgermoell /oellm-longctx-tokenized-superlong-512k-1m-2m-v1 OELLM Superlong Long-Context Tokenized 512K/1M/2M v1 This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows. The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component. Source families: RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.tabulartext-generationn<1K0 likes54 downloads3mo agoHugging Face24m-ric /TRM-modified-datamix-tokenized TRM modified datamix (tokenized) Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model) training, built by running data_io — the HRM-Text data pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three deliberate, documented deviations (below). It is emitted in the V1 tokenized dataset format (a single concatenated token pool + per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.tabulartext-generationn<1K0 likes53 downloads3mo agoHugging Face25hi-todayis-jh /fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes52 downloads7d agoHugging Face26masked-token /solarVQA SolarVQA: A Benchmark for Visual Question Answering in Photovoltaic Defect Inspection Dataset Summary SolarVQA is a structured Visual Question Answering dataset built from expert-annotated electroluminescence (EL) images of silicon solar cells. It contains 130,712 QA pairs across 16,339 images spanning eight complementary question types designed to probe defect existence, counting, type identification, severity, localisation, co-occurrence, and spatial distribution.… See the full description on the dataset page: https://huggingface.co/datasets/masked-token/solarVQA.tabularvisual-question-answering100K<n<1M0 likes48 downloads5mo agoHugging Face27soar-eleuther-i6-hierarchy /metrics-outputs-pcfg-matryoshka-fmt-0000-s2-token-cachetabularn<1K0 likes48 downloads1mo agoHugging Face28open-llm-leaderboard /NucleusAI__nucleus-22B-token-500B-detailsgated Dataset Card for Evaluation run of NucleusAI/nucleus-22B-token-500B Dataset automatically created during the evaluation run of model NucleusAI/nucleus-22B-token-500B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NucleusAI__nucleus-22B-token-500B-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face29ctaxnagomi /deckergui-token-usage-logs Token usage analytics dataset from DeckerGUI ecosystem. Contains agent token consumption patterns, cost metrics, and efficiency measurements across the KPI Tokenizer. Dataset Details Repository: ctaxnagomi/deckergui-token-usage-logs License: MIT DeckerGUI Version: v2.0.0 Created: 2026-08-17 Dataset Schema See metadata.json for the full schema definition. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/deckergui-token-usage-logs.tabularn<1K0 likes43 downloads1mo agoHugging Face30fillay /LWT-2.5B-model-6-tokenizedtabularn<1K0 likes40 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.