datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.SWEUniverse-SizeMatched-45-ReliableTests-20260522
SWEUniverse Size-Matched 45 Repo Reliable Tests
This dataset contains the reliable-test universe rows collected for the 45 repositories in selected_repos_1000_size_matched.tsv.
Contents
data/reliable_tests.jsonl: one strict-stability reliable-test row per completed repo at the 600s timeout bucket. Rows include actual stable_passing, stable_failing, excluded_tests, raw_log_refs, parser metadata, command, commit, and replica counts.
data/snapshot_observations.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/SWEUniverse-SizeMatched-45-ReliableTests-20260522.barexam_256_size_60_test_model
