datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.goldensets
LEGEX Goldensets: Expert-Coded Review-Table Annotations
This repository contains the expert-coded gold annotations for the LEGEX
benchmark of civil-judgment review-table extraction. 1,548 judgments across
19 jurisdictions have been annotated by hand against a shared 14-field schema
covering monetary outcomes, cost allocation, party structure, and industry
classification. Including independent secondary re-annotations, the release
holds 1,974 annotation rows.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/goldensets.economic-narratives-golden-set
Economic Narratives Golden Set
A manually annotated dataset of 500 Russian-language Telegram posts labeled for economic narrative presence, with LLM-generated temporal contexts. This is the evaluation benchmark from the paper on LLM-based economic narrative detection.
Associated Paper
Going Viral: LLM-Based Modeling of Economic Narratives
Dataset Description
The Golden Set was sampled from the Economic Telegram News Corpus to ensure coverage across virality… See the full description on the dataset page: https://huggingface.co/datasets/bruhwalkk/economic-narratives-golden-set.
