datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
movie_reviews_with_context_drift
Dataset Card for reviews_with_drift
Dataset Description
Dataset Summary
This dataset was crafted to be used in our tutorial [Link to the tutorial when ready]. It consists on a large Movie Review Dataset mixed with some reviews from a Hotel Review Dataset. The training/validation set are purely obtained from the Movie Review Dataset while the production set is mixed. Some other features have been added (age, gender, context) as well as a made up timestamp… See the full description on the dataset page: https://huggingface.co/datasets/arize-ai/movie_reviews_with_context_drift.deepseek-1m-context-benchmark
DeepSeek 1M Context Benchmark
This dataset is the publication-safe measurement release for DeepSeek 1M Context Benchmark: Retrieval Accuracy, Latency, and Cost, version v1.0.0. It contains 344 sanitized terminal API records produced by the frozen protocol deepseek-v4-long-context-retrieval-v1.1.0 during a bounded run from 2026-08-06T20:17:02.706Z through 2026-08-07T00:07:44.737Z.
The study compared deepseek-v4-flash and deepseek-v4-pro on deterministic synthetic English… See the full description on the dataset page: https://huggingface.co/datasets/chatdeepai/deepseek-1m-context-benchmark.context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.reality-check-on-context-utilisation
Dataset card for the dataset used in "A Reality Check on Context Utilisation for Retrieval-Augmented Generation"
Dataset Details
This dataset was used for the analysis and plots in the paper "A Reality Check on Context Utilisation for Retrieval-Augmented Generation". More details on the dataset can be found in the paper.
Dataset Description
The dataset contains samples from CounterFact (Meng et al. 2022), ConflictQA (Xie et al. 2024), and DRUID with… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/reality-check-on-context-utilisation.Selective-Context-Gemma3-12B-resultsmodel-context-windows
LLM Context Windows — 206 models
Context-window sizes for 206 ready models served by the Qubax AI API (OpenAI-compatible), exported from the public /v1/models endpoint.
Columns
Column
Description
model_id
API model identifier
model_name
Display name
context_window_tokens
Max context window (tokens)
max_output_tokens
Max output (tokens, where published)
source
Provenance
Notes
License: CC0 1.0 (public domain) — use freely in… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/model-context-windows.Selective-Context-Llama3.1-8B-resultscontext_toxicityhttps://github.com/ipavlopoulos/context_toxicity/
@inproceedings{xenos-etal-2021-context,
title = "Context Sensitivity Estimation in Toxicity Detection",
author = "Xenos, Alexandros and
Pavlopoulos, John and
Androutsopoulos, Ion",
booktitle = "Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021)",
month = aug,
year = "2021",
address = "Online",
publisher = "Association for Computational Linguistics",
url =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/context_toxicity.EmoPillars-Contextless-Processedseo-descriptions-contextGA_Long_Context_Jailbreak_Benchmark
GA Long Context Bench
A benchmark of 1500 multi-turn conversations designed to stress guardrails in long contexts. Each dialog pairs a serialized agent trace with optional prompt injections or per-policy adjudications. Half of the rows embed malicious content deep inside long instructions, enabling evaluation of long-context systems.
Accompanying guardrail releases: GA Guard Core and GA Guard Lite. Check out public benchmarks and results in our blogpost.
[!Note]
Disclaimer: This… See the full description on the dataset page: https://huggingface.co/datasets/GeneralAnalysis/GA_Long_Context_Jailbreak_Benchmark.chatgpt_filtered_sft_traces_context_awareclinical-placebo-contextual-expectation-modulation-mapping-v0.1What this dataset tests
Whether a model can map how context shifts expectation.
Required outputs
expectation_gain_or_penalty_-50_to_50
placebo_amplification_index_0_100
nocebo_risk_index_0_100
Typical failures
ignoring trust and ritual intensity
calling risk disclosure placebo
outputting scores without directional justification
Suggested prompt wrapper
System
You estimate contextual expectation modulation.
User
Baseline expectation{baseline_expectation_0_100}
Clinician… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-placebo-contextual-expectation-modulation-mapping-v0.1.cyberscale-contextual-training
CyberScale Contextual Severity Training Data
Training dataset for the CyberScale contextual severity classifier (Phase 2). Contains 32,000 scenarios combining CVE descriptions with NIS2 sector deployment contexts and cross-border exposure.
Schema
Column
Type
Description
input_text
string
Formatted input: <description> [SEP] sector: <id> cross_border: <bool> score: <float>
label
int
Severity class (0-3)
sector
string
NIS2 sector identifier
cross_border… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-contextual-training.clinical-quad-prior-text-context-window-loss-new-data-summary-extension-hallucination-v0.1What this repo does
This dataset models hallucinated narrative continuation in clinical summaries. It predicts when the interaction between prior text similarity, context window loss, lack of new data, and high summary extension rate indicates that new narrative content has been generated without supporting evidence.
Core quad
prior_text_similarity_index
context_window_loss_index
new_data_presence_index
summary_extension_rate_index
Prediction target
label_hallucinated_continuation
Row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-prior-text-context-window-loss-new-data-summary-extension-hallucination-v0.1.longer-context-hindi-2transparent_context_usage_triale15-context-budget
SignalDepth E15 Context Budget
This is a small prompt-sensitivity benchmark slice for separating two explanations that often get conflated:
the prompt is too short
the task contract is underspecified
The narrow result: on this deterministic Python code-task suite, making sparse prompts longer did not help. Making the task contract explicit did.
Key Result
Condition
Average pass rate
Read
short_sparse
0.25
short and underspecified
long_sparse
0.25
longer… See the full description on the dataset page: https://huggingface.co/datasets/signaldepth/e15-context-budget.Selective-Context-processedpharma-pharmacological-vs-contextual-signal-separation-v0.1What this dataset tests
Whether a system can separate pharmacological effect from contextual modulationin placebo-controlled clinical trial settings.
It treats context as a learnable component of the outcome manifold.
Required outputs
pure_pharmacological_effect
contextual_amplification_factor
nocebo_risk_index
signal_separation_confidence
effect_stability_score
Use case
Second layer of the Placebo/Nocebo Response Disentanglement Matrix.
Improves efficacy estimation by preventing:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/pharma-pharmacological-vs-contextual-signal-separation-v0.1.ai-5node-ctx-buf-lag-cpl-context-overflow-v0.1
What this repo does
This dataset models context overflow cascades in AI agent systems. It detects when context pressure rises, safety buffers weaken due to relaxed truncation and pinning, governance lag delays intervention, and tight coupling through shared summaries propagates constraint loss across workflows, crossing the five-node cascade threshold into an unrecoverable context overflow cascade.
This dataset models a five-node cascade: four interacting instability drivers and one… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-ctx-buf-lag-cpl-context-overflow-v0.1.
