datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DHSA_Long-Data-Collections
DHSA_Long-Data-Collections
A length-bucketed release of togethercomputer/Long-Data-Collections, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight).
Each example is assigned to exactly one length bucket based on token count with meta-llama/Llama-3.1-8B-Instruct (add_special_tokens=True):
Bucket
Token range
lt_8k
[0, 8K)
8k_16k
[8K, 16K)
16k_32k
[16K, 32K)
32k_64k
[32K, 64K)… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_Long-Data-Collections.Long-Data-Collections-booksum-binidxFine-tune Data
BookSum:
BookSum is a dataset for long context summarization. It includes a vast collection of books from various genres, and the task is to generate a coherent and concise summary given a long context from the book. This dataset is designed to test and train models on their ability to understand and summarize long, complex narratives.
to convert to binidx format.
collection_data_clef_checkthat_2026data_set_collectionCollection_of_basic_data
