datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.tokenised_subsetof_erickfmm__red_pajama_es_hq_35tokenizedmultilingual_tokenizer_benchmark
Multilingual Tokenizer Benchmark
More details of each subset like word count, character count, original sources, etc, can be found in the dataset_meta.yaml file in the repository root.
Natural language word count functions
Download spacy models
pip install ntlk spacy pygments underthesea camel-tools
python -m spacy download ko_core_news_sm
python -m spacy download ja_core_news_sm
python -m spacy download zh_core_web_sm
import nltk
nltk.download('punkt_tab')… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/multilingual_tokenizer_benchmark.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.jora_corpus1_FR_tokenized_128kfpabl1-arm-b-fp-tokens-48k
fpabl1-arm-b-fp-tokens-48k
Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp.
Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121).
Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total.
Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks
Source token composition:
fp_en: 1,000,000,000
fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
metrics-outputs-pcfg-matryoshka-layer-03-token-cachedata-32k-200b-tokens
TR-HASH 32K · 200B Token Mixture
Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora/<source>/tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.
Mixture
Source
Weight
Training tokens
DCLM
45%
90B
FineWeb-Edu deduplicated
30%
60B
Stack-Edu
10%
20B… See the full description on the dataset page: https://huggingface.co/datasets/Pacific-i64/data-32k-200b-tokens.metrics-outputs-pcfg-matryoshka-layer-02-token-cachefixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
metrics-outputs-pcfg-matryoshka-layer-01-token-cachemetrics-outputs-pcfg-matryoshka-layer-00-token-cachefixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
metrics-outputs-pcfg-matryoshka-fmt-2400-s1-token-cachemetrics-outputs-pcfg-matryoshka-fmt-2308-s2-token-cachetokenizer-tax
owóorí: the tokenizer tax, measured
Token-count premiums for the 204 languages of FLORES-200 under 12
tokenizers, from the identical 1,012 professionally translated sentences.
The premium is tokens(language) / tokens(English) on the same content, so
it reads directly as a price multiplier for API cost, latency and context
shrinkage.
Interactive explorer: https://kenny0bi.github.io/owoori/
Method, figures, code: https://github.com/Kenny0bi/owoori
Files… See the full description on the dataset page: https://huggingface.co/datasets/kenny0bi/tokenizer-tax.metrics-outputs-pcfg-matryoshka-fmt-2400-s2-token-cachetokenmonster-englishcode-32000-consistent-v1-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
metrics-outputs-pcfg-matryoshka-fmt-2400-s0-token-cachemetrics-outputs-gemma-2-2b-layer-06-token-cacheoellm-longctx-tokenized-superlong-512k-1m-2m-v1
OELLM Superlong Long-Context Tokenized 512K/1M/2M v1
This dataset is a superlong-context continuation-training add-on for extending beyond 256K toward 1M-2M context windows.
The design goal is not simply longer packed text. At 1M-2M, the model needs collection-level continuity and explicit pressure to use very old evidence. This artifact therefore mixes real long structured/natural sources with a small multilingual full-span recall component.
Source families:
RFC Editor… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-superlong-512k-1m-2m-v1.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
solarVQA
SolarVQA: A Benchmark for Visual Question Answering in Photovoltaic Defect Inspection
Dataset Summary
SolarVQA is a structured Visual Question Answering dataset built from
expert-annotated electroluminescence (EL) images of silicon solar cells.
It contains 130,712 QA pairs across 16,339 images spanning eight
complementary question types designed to probe defect existence, counting,
type identification, severity, localisation, co-occurrence, and spatial
distribution.… See the full description on the dataset page: https://huggingface.co/datasets/masked-token/solarVQA.metrics-outputs-pcfg-matryoshka-fmt-0000-s2-token-cacheNucleusAI__nucleus-22B-token-500B-details
Dataset Card for Evaluation run of NucleusAI/nucleus-22B-token-500B
Dataset automatically created during the evaluation run of model NucleusAI/nucleus-22B-token-500B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NucleusAI__nucleus-22B-token-500B-details.deckergui-token-usage-logs
Token usage analytics dataset from DeckerGUI ecosystem. Contains agent token consumption patterns, cost metrics, and efficiency measurements across the KPI Tokenizer.
Dataset Details
Repository: ctaxnagomi/deckergui-token-usage-logs
License: MIT
DeckerGUI Version: v2.0.0
Created: 2026-08-17
Dataset Schema
See metadata.json for the full schema definition.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/deckergui-token-usage-logs.LWT-2.5B-model-6-tokenized
