datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hh-harmless-base-qwen3-8b-margin-dpo-margin-logsultrafeedback-qwen3-8b-margin-dpo-margin-logsfixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts
fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
douvras-quote-margin-reasoning
Douvras Quote and Margin Reasoning v0.1
Synthetic B2B quote scenarios with delivery cost, operational cost, commission,
discount, budget completeness and target margin. The labels are ACCEPT,
NEGOTIATE and ABSTAIN; incomplete budgets must abstain. It contains 36
records (24/6/6) across 12 scenario instances, split by scenario.
This is a calculation protocol, not financial advice. Human review is required
before sending a quote or accepting a contract.
fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts
fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts
er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
fixed-n-rb-offset-marginrl-qwen3-4b-base-polaris53k-offset512-token-mean-1epoch-rollouts
fixed_n_rb_offset_marginrl_Qwen3-4B-Base_polaris53k_offset512_token_mean_1epoch rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
hh-harmless-base-llama3-8b-margin-dpo-margin-logshh-helpful-base-llama3-8b-margin-dpo-margin-logsultrafeedback-llama3-8b-margin-dpo-margin-logshh-helpful-base-qwen3-8b-margin-dpo-margin-logs
