datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imgsdar_webshop_sft_ckpt
sdar_8b_webshop_bs4_eighth_reason_epoch1
This model is a fine-tuned version of /home/hal-yuyangq/Codes/efficient-verl-agent-lyy/temp/SDAR/training/model/SDAR-8B-Chat on the sdar_webshop_bs4_eighth_reason dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The… See the full description on the dataset page: https://huggingface.co/datasets/loongyy/sdar_webshop_sft_ckpt.sdar_4b_webshop_sft_ckpt
sdar_4b_webshop_bs4_eighth_reason_epoch1
This model is a fine-tuned version of /home/hal-yuyangq/Codes/efficient-verl-agent-lyy/SDAR/training/model/SDAR-4B-Chat on the sdar_webshop_bs4_eighth_reason dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The… See the full description on the dataset page: https://huggingface.co/datasets/loongyy/sdar_4b_webshop_sft_ckpt.eigenbench-oct-dpo-vs-introspection
EigenBench OCT: DPO vs Introspection — Scenario-Level Wins
This dataset contains the scenarios on which a DPO-trained persona model (DPO-final)
is judged to be more aligned with a target persona constitution than an Introspection-trained
persona model (Introspection-final), aggregated across multiple judges and orderings.
The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy)
set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.
