datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
radread-public-results
RadRead — public results
Rollout-level results for RadRead, a benchmark of frontier models reading 150
radiographs. Every row is one graded model read: 5 saved rollouts per study
per model, scored by a deterministic grader (no judge model).
A read passes only when every required checklist finding, lesion box (the grader's
IoU / centre / containment test), lexical diagnosis check and action-set membership
check match the reference rubric. No partial credit inside a study;… See the full description on the dataset page: https://huggingface.co/datasets/tirandazdylan/radread-public-results.aidm-dogs-vs-cats-results
aidm-dogs-vs-cats-results
The experiment record of a dogs-vs-cats image-classification study, with CIFAR-10 and CIFAR-10-LT transfer and class-imbalance ablations. This repo holds the run registry, the splits, the report tables and figures, and the per-run predicted probabilities. It holds no images and no model weights; the checkpoints are in the companion model repo.
Generated by scripts/90_publish_hf.py on 2026-09-22 19:17 UTC. Every count, fingerprint and metric below was… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/aidm-dogs-vs-cats-results.
