datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
submission-eval-artifacts
NeurIPS ED 2026 Anonymous Evaluation Artifacts
This dataset repo contains sanitized evaluation artifacts for an anonymous NeurIPS ED 2026 submission. It is metadata-focused: normalized benchmark JSON, selected small paper-facing summaries, reviewer indexes, and manifests.
Checkpoint artifacts are referenced through neurips-ed2026-anon-checkpoints/submission-checkpoints. This dataset repo does not contain model checkpoints or model weights.
Anonymous code artifact:… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed2026-anon-checkpoints/submission-eval-artifacts.forecastgen-artifacts
Forecast-Generalization: raw evaluation outputs across 38 reasoning models
Complete generation-level outputs, per-seed scores and analysis artifacts from a
study of how well benchmark performance forecasts generalization to held-out
reasoning tasks.
Most released evaluations report only aggregate accuracy. This release keeps the
raw per-problem, per-seed generations, so item-level analyses can be redone
without re-running any inference.
What is here
38 models… See the full description on the dataset page: https://huggingface.co/datasets/dvader13/forecastgen-artifacts.magic-video-artifacts
MAGIC-Video — Preprocessing Artifacts
This dataset hosts the exact preprocessing artifacts used in the paper
"Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning"
(MAGIC-Video, arXiv:2605.08271).
Why release these?
The paper's preprocessing pipeline calls LLMs through OpenRouter (translation, OpenIE, semantic
consolidation, narrative chain distillation). Those calls cost money, take hours per subject,
and are non-deterministic — re-running… See the full description on the dataset page: https://huggingface.co/datasets/jiazhengli7/magic-video-artifacts.mhqa-itu-artifacts
MHQA · ITU · Zindi Challenge — Artifacts
DariusTheGeek/mhqa-itu-artifacts · the data + precomputed features that let the code repo reproduce
submission sub_v40 (public LB 0.728509) for the ITU Multilingual Health QA in Low-Resource African
Languages challenge. Code (which pulls this at runtime) lives on GitHub; trained weights are in the model repo
DariusTheGeek/mhqa-itu-adapters.
This is a reproducibility artifact bundle, not a raw dataset. It holds derived features and the… See the full description on the dataset page: https://huggingface.co/datasets/DariusTheGeek/mhqa-itu-artifacts.
