CoolFace
Datasetpublic

jet-ai/pg19-subsample

Jet-Long Evaluation Datasets This repository contains the evaluation data used in the paper Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE. Code: GitHub Repository Dataset Description These datasets are processed versions of standard benchmarks used to evaluate the long-context capabilities of Jet-Long: RULER-500: A dataset for evaluating long-context understanding and retrieval up to 128K context. PG-19-subsample: A subsampled version of… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/pg19-subsample.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes24downloads
Dataset Card

Jet-Long Evaluation Datasets

This repository contains the evaluation data used in the paper Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE.

Dataset Description

These datasets are processed versions of standard benchmarks used to evaluate the long-context capabilities of Jet-Long:

  • —RULER-500: A dataset for evaluating long-context understanding and retrieval up to 128K context.
  • —PG-19-subsample: A subsampled version of the PG-19 dataset used for evaluating long-context language modeling and perplexity.

Citation

If you find Jet-Long or these datasets useful, please cite:

bibtex
@misc{jetlong2026,
  title={Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE},
  author={Haozhan Tang and Zerui Wang and Yuxian Gu and Song Han and Han Cai},
  year={2026},
  eprint={2607.07740},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2607.07740},
}