CoolFace
Datasetpublic

Emulated-Inc/long-context-retrieval-training-pool

Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.

sourceHugging Facecc-by-sa-4.0updated 14d agoView on Hugging Face
1likes128downloads

Emulated-Inc/long-context-retrieval-training-pool · main · files are served by the source, never re-hosted here