datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
environmental_2kgovernance_2ksocial_2kproteus-2k
Proteus-2k
Proteus-2k is a large-scale benchmark table of recent open-weight language models evaluated with the Open LLM Leaderboard v2 pipeline. It was built to extend public leaderboards after freeze dates and to support research on how compute–capability relationships hold up as model families and post-training evolve.
Dataset (full): hlzhang109/proteus-2k — CSV: proteus_2k.csv (~2.4k model checkpoints).
Selected subset: hlzhang109/proteus-selected — CSV: proteus_2k_selected.csv… See the full description on the dataset page: https://huggingface.co/datasets/hlzhang109/proteus-2k.europe-holdout-2k
Europe-Holdout-2k
A 2,000-location street-level geolocation benchmark sampled from the
held-out test fold of the thesis training pipeline. Designed as a
high-variation, contamination-free probe for the published thesis models
(v14 unfreeze_last2, v15, ...) and any external geolocation system.
Locations
2,000
Region
Europe (42 countries)
Source
seed=42 location-level 80/10/10 split of the thesis training corpus
Imagery
Coordinates only — Google Street View bytes… See the full description on the dataset page: https://huggingface.co/datasets/lebfla11/europe-holdout-2k.mnli_test_2Ktesting_2k_data
