TAUR-dev/D-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft
D-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft This evaluation dataset was created as part of the back_to_og_mix__simple_retries__sbon experiment using the SkillFactory experiment management system. Experiment Tracking ๐ View complete experiment details: Experiment Tracker Dataset Evaluation Details {"model": "TAUR-dev/M-back_to_og_mix__simple_retries__sbon-sft", "tasks": ["countdown_2arg", "countdown_3arg"โฆ See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft.
D-EVAL_standardevalv3backtoogmix_simpleretries_sbon-evalsft
This evaluation dataset was created as part of the back_to_og_mix__simple_retries__sbon experiment using the SkillFactory experiment management system.
Experiment Tracking
๐ View complete experiment details: Experiment Tracker Dataset
Evaluation Details
{"model": "TAUR-dev/M-backtoogmixsimpleretries_sbon-sft", "tasks": ["countdown2arg", "countdown3arg", "countdown4arg", "countdown5arg", "countdown6arg", "commonsenseQA", "gsm8k", "longmult2dig", "longmult3dig", "longmult4dig", "longmult5dig"], "annotators": ["greedy"], "splits": ["test"], "dataseturl": "TAUR-dev/D-DATA-canonicaldatasetsplits-v1-71325", "stagename": "evalsft", "uploadtoseparaterepo": true, "mutatepromptforanswertags": true, "maxretries": 1, "requesttimeout": 30, "maxrequestsperminute": 250, "maxinflightrequestslower_multiplier": 1.0}
