CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DinoStackAI /narrativeqa-rag NarrativeQA RAG Dataset for Retrieval-Augmented Generation (RAG) based on NarrativeQA. Structure Subset Splits Description corpus train (default) Wikipedia plot summaries shared across all query splits queries train, dev, test Reading comprehension questions qrels train, dev, test Relevance judgments (query ↔ document) answers train, dev, test Reference answers (longest annotated answer) Dataset statistics Split Queries… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/narrativeqa-rag.tabularquestion-answering100K<n<1M0 likes201 downloads2mo agoHugging Face02CLS-Lab /narrative-gold-annotations Narrative annotation dataset Human annotations for three narrative-analysis tasks — setting, agency, and event relation — over passages sampled from the Dolma corpus. Annotators & anonymization Annotator identities are anonymized. Each task has a single gold adjudicator plus one or more secondary annotators used for double annotation / agreement. Role Meaning gold The adjudicated / primary label for every released instance. annotator_1 Second… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-gold-annotations.tabulartext-classification1K<n<10K1 likes95 downloads2mo agoHugging Face03CLS-Lab /narrative-llm-annotations NarraDolma LLM-Labeled — Distillation Set The intermediate, LLM-labeled dataset that bridges the small human gold set and the full NarraDolma corpus. It contains 5,000 passages sampled from Dolma and labeled by Gemma across all 11 narrative dimensions, stratified by source and topic to preserve the original distribution. These labels are the knowledge-distillation training set used to train NarraBert. Paper: arXiv:2606.19468 Collection: Narratives in LLM Pretraining Data… See the full description on the dataset page: https://huggingface.co/datasets/CLS-Lab/narrative-llm-annotations.tabular10K<n<100K1 likes90 downloads14d agoHugging Face04ClarusC64 /legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does You receive scope billing guidelines time entries fee earner level billing narrative duration and rates flags You decide coherent or incoherent Daily use fee dispute risk scan scope drift scan block billing detection seniority mismatch detection tabulartext-classificationn<1K0 likes48 downloads7mo agoHugging Face05ClarusC64 /legal-billing-narrative-time-entry-coherence-risk-v0.1What this dataset does You receive time entries summary phase and codes invoice narrative totals dup flags client updates You decide coherent or incoherent Daily use bill narrative QC time entry duplication detection dispute risk flagging tabulartext-classificationn<1K0 likes44 downloads7mo agoHugging Face06ClarusC64 /clinical-narrative-coherence-outcome-correlation-mapping-v0.1What this dataset tests Whether narrative coherenceis structurally correlated withclinical outcomes and resilience. Required outputs narrative coherence score outcome alignment score resilience correlation index relapse risk modifier adherence influence signal narrative–outcome relationship Use case Third layer of the Healing Narrative Coherence Corpus. tabulartabular-classificationn<1K0 likes38 downloads8mo agoHugging Face07brandburner /i-claudius-narrative-kg I, Claudius Complete Series Narrative Knowledge Graph Dataset Description This dataset contains a comprehensive narrative knowledge graph extracted from all 13 episodes of the BBC's "I, Claudius" (1976), analyzed using the Fabula V2 pipeline. The graph captures the complex web of Roman imperial politics, family dynamics, and power struggles across the reigns of Augustus, Tiberius, Caligula, and Claudius. Dataset Summary Total Nodes: 10,357 Total Relationships:… See the full description on the dataset page: https://huggingface.co/datasets/brandburner/i-claudius-narrative-kg.tabulargraph-ml10K<n<100K0 likes34 downloads1y agoHugging Face08ClarusC64 /clinical-narrative-image-integrity-v0.2 Clinical Narrative Image Integrity v0.2 What this is A small dataset that tests one question: Can you detect when a clinical narrative-image system is moving toward integrity failure, not just carrying ambiguity? This repo focuses on narrative-image integrity under clinical reasoning pressure. It models a system where: narrative coherence may weaken image alignment may drift interpretive distortion may rise fragmented signal may destabilize representation before overt… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-image-integrity-v0.2.tabulartext-classificationn<1K0 likes33 downloads6mo agoHugging Face09OusiaResearch /narrative-bench Narrative Identity Effect on LLM Reasoning Description {'model': Value('string'), 'biography_level': Value('int64'), 'persona_type': Value('string'), 'task_type': Value('string'), 'accuracy': Value('bool'), 'prompt': Value('string'), 'completion': Value('string'), 'completion_len': Value('int64'), 'lexical_diversity': Value('float64'), 'mattr': Value('float64'), 'sentiment_valence': Value('float64'), 'sentiment_arousal': Value('float64'), 'self_reference_rate':… See the full description on the dataset page: https://huggingface.co/datasets/OusiaResearch/narrative-bench.tabulartext-generation1K<n<10K0 likes31 downloads4mo agoHugging Face10ClarusC64 /clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1 Each row is a trial monthly snapshot. Core quad Data cut timingDatabase lock pressureQuery backlogCSR narrative drift Target label_regulatory_issue_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability. The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.tabulartext-classificationn<1K0 likes29 downloads7mo agoHugging Face11VynFi /vynfi-sar-narratives VynFi SAR Narratives — AML Labels with Transaction Evidence 156 787 banking transactions paired with 156 714 AML labels and case-level SAR (Suspicious Activity Report) narrative text. Designed as a starting point for SAR NLP research and for end-to-end pipelines that go from raw banking activity → AML labels → human-readable narrative. Generated with DataSynth (banking + narrative modules) · GitHub · Companion paper (SSRN). Provenance note. This dataset was last refreshed under… See the full description on the dataset page: https://huggingface.co/datasets/VynFi/vynfi-sar-narratives.tabulartext-generation100K<n<1M0 likes29 downloads5mo agoHugging Face12ClarusC64 /clinical-quad-adherence-dose-ae-efficacy-narrative-collapse-v0.2Clinical Quad Adherence Dose AE Efficacy Narrative Collapse v0.2 What this dataset does It tests whether a system can detect narrative collapse in a clinical decision loop. It forces reasoning across four operational drivers plus the narrative layer. Core quad nodes Adherence stability Dose action AE signal Efficacy signal Narrative node aligned means the story matches the data and governance constraints spin means the story tries to conceal or reframe misalignment What the model… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adherence-dose-ae-efficacy-narrative-collapse-v0.2.tabulartext-classificationn<1K0 likes27 downloads7mo agoHugging Face13jsisonou /narrative-engine-emotion-5c Try the PV Peak/Valley Explorer🔗 PV Radar (Beta) Space: https://huggingface.co/spaces/jsisonou/narrative-engine-pv-radar-betaUse this dataset’s sample files to test: Curve Mode: upload book_curve.scene.csv → Run Text Mode: paste one scene per line → RunYou’ll get pv_pred (per-scene labels), arc_summary (global peak/valley), and score curves.Assistive only; human-in-the-loop. No model weights or training recipes are exposed. ⚠️ This repository is no longer maintained.👉 Please visit the… See the full description on the dataset page: https://huggingface.co/datasets/jsisonou/narrative-engine-emotion-5c.tabularn<1K0 likes26 downloads1y agoHugging Face14ClarusC64 /market-narrative-acceleration-detection-v0.1What this dataset tests Whether a system can detect market narrative accelerationacross heterogeneous sources. This is not sentiment scoring.This is propagation speed and convergence. Required outputs narrative theme acceleration score consensus distance pricing gap state trade signal window Pricing gap states underpriced partially priced priced in neutral Trade signal windows entry early entry fast watch late exit Constraints Do not predict single-asset… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-acceleration-detection-v0.1.tabulartabular-classificationn<1K0 likes23 downloads8mo agoHugging Face15Nachammai41 /remittance-fraud-narratives This dataset is a remastered version prepared using Adaption's Adaptive Data platform. remittance_fraud_narratives This dataset contains prompts designed to generate first-person narratives about financial fraud targeting immigrant communities via cross-border remittance services. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, sender demographics, and language context. The samples currently show null completions, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives.tabularn<1K0 likes20 downloads6mo agoHugging Face16Nachammai41 /itin-fraud-narratives This dataset is a remastered version prepared using Adaption's Adaptive Data platform. itin_fraud_narratives This dataset contains prompt templates designed to generate first-person narratives about financial fraud targeting ITIN holders. Each entry specifies variables such as fraud vector, financial instrument, transaction amount, and language to guide the creation of synthetic victim stories. The samples focus on scenarios involving identity theft, tax fraud, and synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/itin-fraud-narratives.tabular1K<n<10K0 likes20 downloads6mo agoHugging Face17Nachammai41 /remittance-fraud-narratives_with_reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. remittance_fraud_narratives This dataset contains prompts designed to generate first-person narratives about financial transactions, specifically focusing on cross-border remittances within immigrant communities. Each entry specifies details such as the transaction archetype, fraud vector, financial instrument, amount, and sender demographics to guide the creation of realistic scam or legitimate… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives_with_reasoning.tabular1K<n<10K0 likes19 downloads6mo agoHugging Face18ClarusC64 /legal-billing-narrative-task-value-coherence-v0.1What this dataset does You receive billing narrative hours case stage task category value signal duplication signals You decide coherent or incoherent Daily use cost draft review client challenge response write-down triage tabulartext-classificationn<1K0 likes18 downloads7mo agoHugging Face19ClarusC64 /clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1What this repo does This dataset models narrative continuity break in clinical trial summaries. It predicts when the interaction between evidence consistency, endpoint signal strength, claim strength, and certainty language indicates that the written narrative has drifted away from the underlying trial results. Core quad evidence_consistency_index endpoint_signal_strength_index claim_strength_index certainty_language_index Prediction target label_narrative_break Row structure Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1.tabulartext-classificationn<1K0 likes18 downloads7mo agoHugging Face20Nachammai41 /gig-worker-fraud-narratives_with_reasoning This dataset is a remastered version prepared using Adaption's Adaptive Data platform. gig_worker_fraud_narratives This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeovers, hacking, and social engineering. Each prompt specifies details including the fraud vector, financial instrument involved, transaction amount, and sender age to guide the creation of… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives_with_reasoning.tabularn<1K0 likes18 downloads6mo agoHugging Face21mattercalm /narrative_qatabularn<1K0 likes17 downloads2y agoHugging Face22yueq92 /NOAA_narratives_NERtabularn<1K0 likes17 downloads2y agoHugging Face23ClarusC64 /clinical-narrative-clinical-timeline-alignment-v0.1What this dataset tests Whether a system can alignpatient-reported narrativeswith objective clinical timelines. Required outputs alignment score narrative time shift omitted events overemphasized events narrative anchors misalignment risk band Use case First layer of the Healing Narrative Coherence Corpus. tabulartabular-classificationn<1K0 likes17 downloads8mo agoHugging Face24ClarusC64 /market-narrative-fragility-and-consensus-peak-v0.1What this dataset tests Whether a system can detectwhen a market narrative reaches consensus saturationand begins to destabilize. This is late-stage narrative detection. Required outputs narrative theme consensus density score fragility score divergence signals exit risk window Exit windows exit_prepare exit_watch exit_fast Constraints Do not predict price direction.Detect narrative saturation and fragility. Evaluation focus High scores requireclear narrative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-fragility-and-consensus-peak-v0.1.tabulartabular-classificationn<1K0 likes17 downloads8mo agoHugging Face25ClarusC64 /clinical-quad-adherence-dose-changes-adverse-events-efficacy-narrative-v0.2Clinical Quad Adherence Dose Changes Adverse Events Efficacy Narrative v0.2 What this dataset does It tests whether a model can classify coherence versus collapse in a clinical decision loop. The loop couples four operational nodes plus narrative behavior. Quad nodes adherence dose changes adverse events efficacy signal Narrative node aligned means the story matches data and constraints spin means the story tries to justify a decision that the data does not support Labels 0… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adherence-dose-changes-adverse-events-efficacy-narrative-v0.2.tabulartext-classificationn<1K0 likes17 downloads7mo agoHugging Face26Nachammai41 /gig-worker-fraud-narratives This dataset is a remastered version prepared using Adaption's Adaptive Data platform. gig_worker_fraud_narratives This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeover, hacking, and social engineering. Each prompt specifies details like the fraud vector, financial instrument, transaction amount, and sender age to guide the creation of realistic scam… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives.tabular1K<n<10K0 likes15 downloads6mo agoHugging Face27Nachammai41 /unbanked-fraud-narratives This dataset is a remastered version prepared using Adaption's Adaptive Data platform. unbanked_fraud_narratives This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in legitimate or fraudulent financial transactions. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, and community context like payday loans or prepaid cards. The completions are currently empty, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/unbanked-fraud-narratives.tabular1K<n<10K0 likes14 downloads6mo agoHugging Face28bruhwalkk /economic-narratives-golden-set Economic Narratives Golden Set A manually annotated dataset of 500 Russian-language Telegram posts labeled for economic narrative presence, with LLM-generated temporal contexts. This is the evaluation benchmark from the paper on LLM-based economic narrative detection. Associated Paper Going Viral: LLM-Based Modeling of Economic Narratives Dataset Description The Golden Set was sampled from the Economic Telegram News Corpus to ensure coverage across virality… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-narratives-golden-set.tabulartext-classificationn<1K0 likes13 downloads5mo agoHugging Face29ClarusC64 /clinical-narrative-implicit-normalization-bias-v0.4 Implicit Normalization Bias Clinical Narrative Integrity v0.4 Purpose This dataset tests whether a model: Avoids assuming normality when data is missing Resists default reassurance Preserves honest narrative boundaries Treats “normal” as a claim, not a default You are measuring baseline discipline. Why this dataset exists Clinical notes often omit information. A failure mode distinct from hallucinated negatives is more subtle: Turning… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-implicit-normalization-bias-v0.4.tabularn<1K0 likes11 downloads8mo agoHugging Face30ClarusC64 /clinical-narrative-negative-evidence-handling-v0.3 Negative Evidence Handling Clinical Narrative Integrity v0.3 Purpose This dataset tests whether a model can: Distinguish absence of documentation from true negative findings Avoid inventing exclusions Preserve epistemic boundaries Maintain honest clinical narrative structure You are measuring restraint, not fluency. Why this matters Clinical documentation is incomplete by default. A safe system must: Say less when less is known Avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-narrative-negative-evidence-handling-v0.3.tabularn<1K0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.