datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
officeqa-checkpoint-eval-data
Checkpoint evaluation plot data
Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures.
No model execution, grading, publication, or source-result changes were performed to make this export.
Contents
checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds.
pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.OfficeFatigue
OfficeFatigue Signal-Only Release
This initial release contains participant-level processed signal arrays for OfficeFatigue. It includes 25 participants (p01-p25) and two subsets per participant:
OFS_signal.npy: short controlled desk-work subset.
OFL_signal.npy: long natural office subset.
Fatigue labels, benchmark splits, and diagnostic/oracle-corrected labels are not included in this signal-only release. Labels will be released separately in a later version.
File… See the full description on the dataset page: https://huggingface.co/datasets/Officefatigue/OfficeFatigue.
