di2ox3/prefill-dataset
Prefill Dataset Long-context tokenized corpus for benchmarking LLM prefill computation with Qwen3-8B. Contains ~10M tokens of copyright-free English text pre-tokenized with character offset mappings for fast position lookup. Dataset Structure Files File Description Rows data/documents.parquet English documents with token IDs and char offsets ~100-500 data/tasks.parquet QA, translation, and retrieval tasks ~1K-5K… See the full description on the dataset page: https://huggingface.co/datasets/di2ox3/prefill-dataset.
This repository belongs to di2ox3 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
