datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.uk-dale-haystack
UK-DALE-Haystack
A controlled additive-needle benchmark for long-context time-series language
models built on top of UK-DALE (Kelly & Knottenbelt, 2015), the canonical
UK domestic appliance-level + whole-house power demand dataset.
Each sample is a 6-second-sampled mains active-power trace with one or more
real per-appliance bouts inserted at known locations. A QA prompt asks the
model to detect, count, localize, order, or reason about those bouts across
five context lengths from 15… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/uk-dale-haystack.capture24-ts-haystack-fixed-needle
Capture24 TS-Haystack — Fixed Needle Length
Long-context retrieval / reasoning benchmark over Capture24 wrist-worn
accelerometer recordings, used in Recursive Agents are Effective Time Series
Reasoners (ARTS-RLM).
This repository supersedes
nz00shuuuu/capture24-ts-haystack-cot
for the paper's main capture24 experiments. Differences:
Fixed (absolute-ms) needle length of 3–10 s across every context length
instead of needles that scale with context. With a 7200 s haystack the
needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.urbansound-haystack
Urban-Sound-Haystack
A long-context urban-audio QA benchmark across 10 task types and
4 context lengths (100 s, 15 min, 30 min, 1 h). Each soundscape is
synthesised by Scaper from
UrbanSound8K foreground events over TUT acoustic-scene backgrounds, sampled
at 16 kHz mono PCM_32. Two of the ten tasks
(anomaly_detection, anomaly_localization) draw from a parallel pool
where every soundscape contains exactly one out-of-vocabulary event from
ESC-50 (glass_breaking or crying_baby).
This… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/urbansound-haystack.
