datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.economic-narratives-golden-set
Economic Narratives Golden Set
A manually annotated dataset of 500 Russian-language Telegram posts labeled for economic narrative presence, with LLM-generated temporal contexts. This is the evaluation benchmark from the paper on LLM-based economic narrative detection.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The Golden Set was sampled from the Economic Telegram News Corpus to ensure coverage across virality… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-narratives-golden-set.
