Glow-AI/pco32_tabular_data
pco32 Throughput of distributed LLM training as a function of the parallelism configuration, on 32 GPUs across 8 hosts. Tabular benchmark for the bolt problem pco32: the 379 rows are the full candidate set. Objective throughput_mean, to maximise. NaN where the run OOMed, since no throughput is observed at all -- a hidden (crash) constraint, not a bad value. Constraint ran_successfully. 339 of 379 configurations run; the rest exhaust GPU memory. Feasibility is only learnable by… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/pco32_tabular_data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face