datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ltaf-haystack-fixedcapture24-ts-haystack-cotltaf-haystackuk-dale-haystack
UK-DALE-Haystack
A controlled additive-needle benchmark for long-context time-series language
models built on top of UK-DALE (Kelly & Knottenbelt, 2015), the canonical
UK domestic appliance-level + whole-house power demand dataset.
Each sample is a 6-second-sampled mains active-power trace with one or more
real per-appliance bouts inserted at known locations. A QA prompt asks the
model to detect, count, localize, order, or reason about those bouts across
five context lengths from 15… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/uk-dale-haystack.sleep_psg_ts_haystackHaystackID_MQP_2025-2026capture24-ts-haystack-fixed-needle
Capture24 TS-Haystack — Fixed Needle Length
Long-context retrieval / reasoning benchmark over Capture24 wrist-worn
accelerometer recordings, used in Recursive Agents are Effective Time Series
Reasoners (ARTS-RLM).
This repository supersedes
nz00shuuuu/capture24-ts-haystack-cot
for the paper's main capture24 experiments. Differences:
Fixed (absolute-ms) needle length of 3–10 s across every context length
instead of needles that scale with context. With a 7200 s haystack the
needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.summary-of-a-haystack
Dataset Card for SummHay
This repository contains the data for the experiments in the SummHay paper.
Accessing the Data
We publicly release the 10 Haystacks (5 in conversational domain, 5 in the news domain). Each example follows the below format:
{
"topic_id": "ObjectId()",
"topic": "",
"topic_metadata": {"participants": []}, // can be domain specific
"subtopics": [
{
"subtopic_id": "ObjectId()",
"subtopic_name": ""… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/summary-of-a-haystack.anti-haystack
Dataset Card for "anti-haystack"
This dataset contains samples that resemble the "Needle in a haystack" pressure testing. It can be helpful if you want to make your LLM better at finding/locating short facts from long documents.
Data Structure
Each sample has the following fields:
document: A long and noisy reference document which can be a story, code, book, or manual in both English and Chinese (10%).
question: A question generated with GPT-4. The answer can always be… See the full description on the dataset page: https://huggingface.co/datasets/wenbopan/anti-haystack.medrag-pubmed-chunk-with-embeddingshaystack-pipelinesneedle-in-a-haystack-biographies-v2haystack-pipelines-v2urbansound-haystack
Urban-Sound-Haystack
A long-context urban-audio QA benchmark across 10 task types and
4 context lengths (100 s, 15 min, 30 min, 1 h). Each soundscape is
synthesised by Scaper from
UrbanSound8K foreground events over TUT acoustic-scene backgrounds, sampled
at 16 kHz mono PCM_32. Two of the ten tasks
(anomaly_detection, anomaly_localization) draw from a parallel pool
where every soundscape contains exactly one out-of-vocabulary event from
ESC-50 (glass_breaking or crying_baby).
This… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/urbansound-haystack.qwen3_0.6b-task738_augmented_needle_in_a_haystack_Mar16-1507_blendedqwen3_0.6b_needle_in_a_haystack_Mar17-1948_blendedc1_haystackmultineedle-3needles-haystack-datasetc3_haystacktask738_augmented_needle_in_a_haystack_Mar16-1507haystack-pipelines-v3needle-in-a-haystack-biographies-v0qwen3_0.6b-rlvr_Feb24-1812_datamix_top10_augmented_needle_in_a_haystack_Feb25-0119qwen3_0.6b-task738_augmented_needle_in_a_haystack_Mar16-1507qwen3_0.6b_needle_in_a_haystack_Mar17-1948brick-haystack-tagsmicro_top2_augmented_needle_in_a_haystack_Mar17-1948haystack-pipelines-no-paramsgpt4o-multineedle-3needles-haystack-dataset_iter2c2_haystack
