datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iterative-dpo-data-for-SimPO-iter2
iterative-dpo-data-for-SimPO-iter2
概要
合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成
5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施
1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置
全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.SFT-Dataset-For-Self-Taught-Evaluators-iter1iterative-dpo-data-for-ORPO-iter3
iterative-dpo-data-for-ORPO-iter3
概要
合成instructionデータであるAratako/Self-Instruct-Qwen2.5-72B-Instruct-60kを元に以下のような手順で作成した日本語Preferenceデータセットです。
開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter2を用いて、temperature=1で回答を5回生成
5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施
1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置
全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外
ライセンス
本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。
META LLAMA 3.1 COMMUNITY… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-ORPO-iter3.debate-iter2-rescored
Debate Iter2 - Rescored Dataset
Training data for debate model GRPO fine-tuning. Contains multi-trial responses scored by Claude Sonnet.
Dataset Configs
Config
File
Rows
Description
flat (default)
rescored_flat.parquet
6,419
One row per call, 4 responses per row
expanded
rescored_samples.parquet
23,403
One row per response
group_a_with_logprobs
group_a_rescored_with_logps.parquet
2,055
Group A with precomputed logprobs
Flat Format Columns… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-iter2-rescored.ipda-grpo-dataset-iter3-feb-12
IPDA GRPO Dataset — Iteration 3 (Feb 12, 2026)
GRPO (Group Relative Policy Optimization) training dataset for IPDA (International Public Debate Association) debate speech generation.
Dataset Structure
2,988 unique prompts | 11,425 scored trials | Score avg: 0.700 (0-1 scale)
Each row represents a unique debate pipeline prompt with up to 6 trial responses:
Column
Description
prompt_hash
SHA256[:16] of prompt text
prompt
Full pipeline prompt
speech_type
AC… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-grpo-dataset-iter3-feb-12.
