datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
holobench
HoloBench (Holistic Reasoning Benchmark)
HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts.
Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.MEGA-cleaned-prompts
Cleaned Prompts Mega Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A comprehensive collection of 2.7 million cleaned English prompts, meticulously processed for training advanced language models and AI systems.
📊 Dataset Statistics
Metric
Value
Total Rows
2,689… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/MEGA-cleaned-prompts.
