64k-context
long-context-attention-labels-64k-128k
Long-Context Attention Labels (64K & 128K)
Attention-based document labels for long-context training data selection.
Overview
Each document is labeled with attention lookback metrics computed by running it through a model and measuring how far back each token attends (via top-k=10 head-averaged attention distances).
Run
Model
Context
Records
olmo3_64k
OLMo-3-1025-7B (stage2)
65,536 tokens
~5,000
olmo3_128k
OLMo-3-1025-7B (stage2)
131,072 tokens
~4,600… See the full description on the dataset page: https://huggingface.co/datasets/KevinDavidHayes/long-context-attention-labels-64k-128k.smoltalk2_LongAlign_64k_context_french_no_think
Description
This is the SFT/LongAlign_64k_context_lang_annotated_lang_6_no_think subset of HuggingFaceTB/smoltalk2, a reasoning dataset, which we have filtered to keep only the French data.In practice, the data contained an English prompt, then French data to be analyzed in relation to the prompt, and finally the model response. We therefore kept the French part, but also translated the French prompt and response.
slimpajama-long-context-pythia-64k
