datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-tokenized
FineWeb Tokenized
> 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer
What is it?
This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards.
By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.seamless-align-enA-jaA.tokenized.encodecfineweb-tokenized-fake
What is it?
It's similar to anisolai/fineweb-tokenized but fake.
I don't understand why I did that :)
WARNING:
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN SOMEONE.
WHY ARE YOU DOWNLOADING IT? YOU COULD LOSE MILLIONS OF DOLLARS IF YOU USE IT TO TRAIN… See the full description on the dataset page: https://huggingface.co/datasets/mondk/fineweb-tokenized-fake.seamless-align-enA-hiA.tokenized.encodecseamless-align-enA-esA.tokenized.encodecseamless-align-enA-viA.tokenized.encodecseamless-align-deA-enA.tokenized.encodeccarbon-tokenized-corpus
Dataset Summary
AINovice2005/carbon-tokenized-corpus is the tokenized sample of AINovice2005/carbon-cpu-enriched-sequences-sampled .
Schema
The current dataset contains the following fields:
Field
Type
Description
record_id
string
Source/reference sequence identifier
start
int64
Start coordinate of the sequence interval
end
int64
End coordinate of the sequence interval
token_ids
list
Integer token IDs produced by the tokenizer
token_mask
list… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-tokenized-corpus.0399-tv-valid-clean-sft-tokenized-llmjp4-8bexp-pool-repository-code-dolma2-tokenized
Locus EXP Repository Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tokenized_sampleseamless-align-enA-koA.tokenized.encodecseamless-align-enA-frA.tokenized.encodecIchigo-instruction-tokenized-v0.2sib200-xlmr-tokenized
SIB-200 Tokenized by XLM-R Large
This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352.
The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper,
ensuring consistent and reproducible tokenization across all languages and experimental settings.
The dataset is organized as a multi-config Hugging… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.mixed-pretraining-tokenizedbinary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.finemath-4plus-tokenizedtokenized_fineweb_edu_10b_combinedseamless-align-enA-zhA.tokenized.encodectokenized_wikipedia_20220301.en_train_512
Tokenized English Wikipedia Dataset
Dataset Description
This dataset contains tokenized chunks of text from the English Wikipedia dump of March 1, 2022. Each entry in the dataset represents a chunk of text from Wikipedia, with information about which document and position within the document it comes from.
Dataset Creation
Source Dataset: Wikipedia (20220301.en)
Tokenizer: BERT base uncased
Chunk Size: 512 tokens (including special tokens)… See the full description on the dataset page: https://huggingface.co/datasets/TemryL/tokenized_wikipedia_20220301.en_train_512.HuggingFaceFW-fineweb-sample-10BT-tokenizedIchigo-instruction-tokenized-v0.2-cleanMaCoCu-sl-tokenized
Dataset Card for MaCoCu-sl Multi-Tokenized
Dataset Description:
This dataset provides a pre-tokenized version of the Slovene web corpus MaCoCu.
It includes the original text data and metadata from MaCoCu-sl, augmented with token IDs and token counts generated by several
popular large language model tokenizers. The goal is to facilitate research and experimentation by providing ready-to-use tokenized data,
saving computational resources during repeated setups.
Licensing and… See the full description on the dataset page: https://huggingface.co/datasets/SemantikaEU/MaCoCu-sl-tokenized.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.HuggingFaceFW-fineweb-edu-sample-10BT-tokenizedleader-training-tokenized-fixed-summary-16k-fullHuggingFaceTB-finemath-finemath-4plus-tokenizedexp-pool-academic-dolma2-tokenized
Locus EXP Academic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.Ichigo-instruction-tokenized-v0.1
