CoolFace
Datasetpublic

di2ox3/prefill-dataset

Prefill Dataset Long-context tokenized corpus for benchmarking LLM prefill computation with Qwen3-8B. Contains ~10M tokens of copyright-free English text pre-tokenized with character offset mappings for fast position lookup. Dataset Structure Files File Description Rows data/documents.parquet English documents with token IDs and char offsets ~100-500 data/tasks.parquet QA, translation, and retrieval tasks ~1K-5K… See the full description on the dataset page: https://huggingface.co/datasets/di2ox3/prefill-dataset.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes26downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
di2ox3/prefill-dataset · CoolFace