CoolFace
Datasetpublic

anonymous-submission-0001/anonymous-submission-01

NarrativeBench NarrativeBench evaluates whether LLMs cite and cover diverse multilingual narrative cells in contested geopolitical information-seeking tasks. Tables data/documents.*: 588 source documents with public narrative IDs and translations. data/narratives.*: 97 public narrative cells. data/questions.*: 213 questions across 9 prompt languages; Reddit URLs are not stored here. data/reddit_sources.*: Reddit source URLs linked to released question IDs.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-0001/anonymous-submission-01.

sourceHugging Facecc-by-nc-sa-4.0updated 22d agoView on Hugging Face
0likes368downloads
Dataset Card

NarrativeBench

NarrativeBench evaluates whether LLMs cite and cover diverse multilingual narrative cells in contested geopolitical information-seeking tasks.

Tables

  • —data/documents.*: 588 source documents with public narrative IDs and translations.
  • —data/narratives.*: 97 public narrative cells.
  • —data/questions.*: 213 questions across 9 prompt languages; Reddit URLs are not stored here.
  • —data/reddit_sources.*: Reddit source URLs linked to released question IDs.
  • —data/eval_2k/*: 1,720 evaluation packs and citation targets used for final scoring.
  • —results/scored_generations.*: 20,052 retained scored generations from 12 model runs.
  • —results/capability_summary_by_model.*: model-level capability table used in the paper.

Notes

Question labels are the eval-time labels used in the final echo-chamber analysis. Gold targets are narrative-ID based: NI keeps exactly the target narrative, NE removes exactly the target narrative, and MNI removes the baseline narrative already represented in the prior answer. context_doc_meta.country and target metadata use the same public perspective/country labels as data/documents.*. Public narrative cells are internally consistent by conflict, country, language, and target phrase; one collided Tigray source ID is split into its China, France, and United Kingdom cells. In eval.jsonl, per-document metadata fields are objects keyed by document ID; in eval.parquet, they are compact lists of records carrying an explicit doc_id.