haystack
embeddinggemma-300m-haystack-contrastive-better-thinembeddinggemma-300m-haystack-contrastive-very-thinembeddinggemma-300m-haystack-contrastive-very-thickembeddinggemma-300m-haystack-contrastive-thin-fixedset_date_1_bert-base-uncased_finetuned_with_haystackts_haystack_itformer_llamaGliner_haystackfalcon7b-ft-haystack
Datasets
All datasets matching “haystack”document-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.document-haystack-10pages
Dataset Card for document-haystack-10pages
This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/document-haystack-10pages")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.HaystackCraft@article{li2025haystack,
title={Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation},
author={Mufei Li and Dongqi Fu and Limei Wang and Si Zhang and Hanqing Zeng and Kaan Sancak and Ruizhong Qiu and Haoyu Wang and Xiaoxin He and Xavier Bresson and Yinglong Xia and Chonglin Sun and Pan Li},
journal={arXiv preprint arXiv:2510.07414},
year={2025}
}
ltaf-haystack-fixedcapture24-ts-haystack-cotuk-dale-haystack
UK-DALE-Haystack
A controlled additive-needle benchmark for long-context time-series language
models built on top of UK-DALE (Kelly & Knottenbelt, 2015), the canonical
UK domestic appliance-level + whole-house power demand dataset.
Each sample is a 6-second-sampled mains active-power trace with one or more
real per-appliance bouts inserted at known locations. A QA prompt asks the
model to detect, count, localize, order, or reason about those bouts across
five context lengths from 15… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/uk-dale-haystack.
