datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
narrativeqa-rag
NarrativeQA RAG
Dataset for Retrieval-Augmented Generation (RAG) based on NarrativeQA.
Structure
Subset
Splits
Description
corpus
train (default)
Wikipedia plot summaries shared across all query splits
queries
train, dev, test
Reading comprehension questions
qrels
train, dev, test
Relevance judgments (query ↔ document)
answers
train, dev, test
Reference answers (longest annotated answer)
Dataset statistics
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/narrativeqa-rag.narrative-gold-annotations
Narrative annotation dataset
Human annotations for three narrative-analysis tasks — setting, agency,
and event relation — over passages sampled from the Dolma corpus.
Annotators & anonymization
Annotator identities are anonymized. Each task has a single gold adjudicator
plus one or more secondary annotators used for double annotation / agreement.
Role
Meaning
gold
The adjudicated / primary label for every released instance.
annotator_1
Second… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-gold-annotations.narrative-llm-annotations
NarraDolma LLM-Labeled — Distillation Set
The intermediate, LLM-labeled dataset that bridges the small human gold set and the
full NarraDolma corpus. It contains 5,000 passages sampled from
Dolma and labeled by Gemma across
all 11 narrative dimensions, stratified by source and topic to preserve the original
distribution. These labels are the knowledge-distillation training set used to
train NarraBert.
Paper: arXiv:2606.19468
Collection: Narratives in LLM Pretraining Data… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-llm-annotations.legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
legal-billing-narrative-time-entry-coherence-risk-v0.1What this dataset does
You receive
time entries summary
phase and codes
invoice narrative
totals
dup flags
client updates
You decide
coherent
or
incoherent
Daily use
bill narrative QC
time entry duplication detection
dispute risk flagging
clinical-narrative-coherence-outcome-correlation-mapping-v0.1What this dataset tests
Whether narrative coherenceis structurally correlated withclinical outcomes and resilience.
Required outputs
narrative coherence score
outcome alignment score
resilience correlation index
relapse risk modifier
adherence influence signal
narrative–outcome relationship
Use case
Third layer of the Healing Narrative Coherence Corpus.
i-claudius-narrative-kg
I, Claudius Complete Series Narrative Knowledge Graph
Dataset Description
This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius.
Dataset Summary
Total Nodes: 10,357
Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.clinical-narrative-image-integrity-v0.2
Clinical Narrative Image Integrity v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical narrative-image system is moving toward integrity failure, not just carrying ambiguity?
This repo focuses on narrative-image integrity under clinical reasoning pressure.
It models a system where:
narrative coherence may weaken
image alignment may drift
interpretive distortion may rise
fragmented signal may destabilize representation before overt… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-image-integrity-v0.2.narrative-bench
Narrative Identity Effect on LLM Reasoning
Description
{'model': Value('string'), 'biography_level': Value('int64'), 'persona_type': Value('string'), 'task_type': Value('string'), 'accuracy': Value('bool'), 'prompt': Value('string'), 'completion': Value('string'), 'completion_len': Value('int64'), 'lexical_diversity': Value('float64'), 'mattr': Value('float64'), 'sentiment_valence': Value('float64'), 'sentiment_arousal': Value('float64'), 'self_reference_rate':… See the full description on the dataset page: https://huggingface.co/datasets/OusiaResearch/narrative-bench.clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1
Each row is a trial monthly snapshot.
Core quad
Data cut timingDatabase lock pressureQuery backlogCSR narrative drift
Target
label_regulatory_issue_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.vynfi-sar-narratives
VynFi SAR Narratives — AML Labels with Transaction Evidence
156 787 banking transactions paired with 156 714 AML labels and
case-level SAR (Suspicious Activity Report) narrative text.
Designed as a starting point for SAR NLP research and for
end-to-end pipelines that go from raw banking activity → AML
labels → human-readable narrative.
Generated with DataSynth (banking + narrative modules) ·
GitHub ·
Companion paper (SSRN).
Provenance note. This dataset was last refreshed under… See the full description on the dataset page: https://huggingface.co/datasets/VynFi/vynfi-sar-narratives.clinical-quad-adherence-dose-ae-efficacy-narrative-collapse-v0.2Clinical Quad Adherence Dose AE Efficacy Narrative Collapse v0.2
What this dataset does
It tests whether a system can detect narrative collapse in a clinical decision loop.
It forces reasoning across four operational drivers plus the narrative layer.
Core quad nodes
Adherence stability
Dose action
AE signal
Efficacy signal
Narrative node
aligned means the story matches the data and governance constraints
spin means the story tries to conceal or reframe misalignment
What the model… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adherence-dose-ae-efficacy-narrative-collapse-v0.2.narrative-engine-emotion-5c
Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test:
Curve Mode: upload book_curve.scene.csv → Run
Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed.
⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/narrative-engine-emotion-5c.market-narrative-acceleration-detection-v0.1What this dataset tests
Whether a system can detect market narrative accelerationacross heterogeneous sources.
This is not sentiment scoring.This is propagation speed and convergence.
Required outputs
narrative theme
acceleration score
consensus distance
pricing gap state
trade signal window
Pricing gap states
underpriced
partially priced
priced in
neutral
Trade signal windows
entry early
entry fast
watch
late
exit
Constraints
Do not predict single-asset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-acceleration-detection-v0.1.remittance-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
remittance_fraud_narratives
This dataset contains prompts designed to generate first-person narratives about financial fraud targeting immigrant communities via cross-border remittance services. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, sender demographics, and language context. The samples currently show null completions, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives.itin-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
itin_fraud_narratives
This dataset contains prompt templates designed to generate first-person narratives about financial fraud targeting ITIN holders. Each entry specifies variables such as fraud vector, financial instrument, transaction amount, and language to guide the creation of synthetic victim stories. The samples focus on scenarios involving identity theft, tax fraud, and synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/itin-fraud-narratives.remittance-fraud-narratives_with_reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
remittance_fraud_narratives
This dataset contains prompts designed to generate first-person narratives about financial transactions, specifically focusing on cross-border remittances within immigrant communities. Each entry specifies details such as the transaction archetype, fraud vector, financial instrument, amount, and sender demographics to guide the creation of realistic scam or legitimate… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives_with_reasoning.legal-billing-narrative-task-value-coherence-v0.1What this dataset does
You receive
billing narrative
hours
case stage
task category
value signal
duplication signals
You decide
coherent
or
incoherent
Daily use
cost draft review
client challenge response
write-down triage
clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1What this repo does
This dataset models narrative continuity break in clinical trial summaries. It predicts when the interaction between evidence consistency, endpoint signal strength, claim strength, and certainty language indicates that the written narrative has drifted away from the underlying trial results.
Core quad
evidence_consistency_index
endpoint_signal_strength_index
claim_strength_index
certainty_language_index
Prediction target
label_narrative_break
Row structure
Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1.gig-worker-fraud-narratives_with_reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gig_worker_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeovers, hacking, and social engineering. Each prompt specifies details including the fraud vector, financial instrument involved, transaction amount, and sender age to guide the creation of… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives_with_reasoning.narrative_qaNOAA_narratives_NERclinical-narrative-clinical-timeline-alignment-v0.1What this dataset tests
Whether a system can alignpatient-reported narrativeswith objective clinical timelines.
Required outputs
alignment score
narrative time shift
omitted events
overemphasized events
narrative anchors
misalignment risk band
Use case
First layer of the Healing Narrative Coherence Corpus.
market-narrative-fragility-and-consensus-peak-v0.1What this dataset tests
Whether a system can detectwhen a market narrative reaches consensus saturationand begins to destabilize.
This is late-stage narrative detection.
Required outputs
narrative theme
consensus density score
fragility score
divergence signals
exit risk window
Exit windows
exit_prepare
exit_watch
exit_fast
Constraints
Do not predict price direction.Detect narrative saturation and fragility.
Evaluation focus
High scores requireclear narrative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-fragility-and-consensus-peak-v0.1.clinical-quad-adherence-dose-changes-adverse-events-efficacy-narrative-v0.2Clinical Quad Adherence Dose Changes Adverse Events Efficacy Narrative v0.2
What this dataset does
It tests whether a model can classify coherence versus collapse in a clinical decision loop.
The loop couples four operational nodes plus narrative behavior.
Quad nodes
adherence
dose changes
adverse events
efficacy signal
Narrative node
aligned means the story matches data and constraints
spin means the story tries to justify a decision that the data does not support
Labels
0… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adherence-dose-changes-adverse-events-efficacy-narrative-v0.2.gig-worker-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gig_worker_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeover, hacking, and social engineering. Each prompt specifies details like the fraud vector, financial instrument, transaction amount, and sender age to guide the creation of realistic scam… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives.unbanked-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
unbanked_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in legitimate or fraudulent financial transactions. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, and community context like payday loans or prepaid cards. The completions are currently empty, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/unbanked-fraud-narratives.economic-narratives-golden-set
Economic Narratives Golden Set
A manually annotated dataset of 500 Russian-language Telegram posts labeled for economic narrative presence, with LLM-generated temporal contexts. This is the evaluation benchmark from the paper on LLM-based economic narrative detection.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The Golden Set was sampled from the Economic Telegram News Corpus to ensure coverage across virality… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-narratives-golden-set.clinical-narrative-implicit-normalization-bias-v0.4
Implicit Normalization Bias
Clinical Narrative Integrity v0.4
Purpose
This dataset tests whether a model:
Avoids assuming normality when data is missing
Resists default reassurance
Preserves honest narrative boundaries
Treats “normal” as a claim, not a default
You are measuring baseline discipline.
Why this dataset exists
Clinical notes often omit information.
A failure mode distinct from hallucinated negatives is more subtle:
Turning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-implicit-normalization-bias-v0.4.clinical-narrative-negative-evidence-handling-v0.3
Negative Evidence Handling
Clinical Narrative Integrity v0.3
Purpose
This dataset tests whether a model can:
Distinguish absence of documentation from true negative findings
Avoid inventing exclusions
Preserve epistemic boundaries
Maintain honest clinical narrative structure
You are measuring restraint, not fluency.
Why this matters
Clinical documentation is incomplete by default.
A safe system must:
Say less when less is known
Avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-negative-evidence-handling-v0.3.
