CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aratako /iterative-dpo-data-for-SimPO-iter2 iterative-dpo-data-for-SimPO-iter2 概要 合成instructionデータであるAratako/Magpie-Tanuki-Instruction-Selected-Evolved-26.5kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter1を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-SimPO-iter2.tabulartext-generation10K<n<100K1 likes41 downloads2y agoHugging Face02Aratako /SFT-Dataset-For-Self-Taught-Evaluators-iter1tabulartext-generation10K<n<100K1 likes37 downloads2y agoHugging Face03Aratako /iterative-dpo-data-for-ORPO-iter3 iterative-dpo-data-for-ORPO-iter3 概要 合成instructionデータであるAratako/Self-Instruct-Qwen2.5-72B-Instruct-60kを元に以下のような手順で作成した日本語Preferenceデータセットです。 開発途中のモデルであるAratako/Llama-Gemma-2-27b-CPO_SimPO-iter2を用いて、temperature=1で回答を5回生成 5個の回答それぞれに対して、Qwen/Qwen2.5-72B-Instruct-GPTQ-Int8を用いて0~5点のスコア付けを実施 1つのinstructionに対する5個の回答について、最もスコアが高いものをchosenに、低いものをrejectedに配置 全て同じスコアの場合や、最も良いスコアが2点以下の場合は除外 ライセンス 本データセットは回答の作成に利用したモデルの関係で以下のライセンスの影響を受けます。 META LLAMA 3.1 COMMUNITY… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/iterative-dpo-data-for-ORPO-iter3.tabulartext-generation10K<n<100K3 likes26 downloads2y agoHugging Face04debaterhub /debate-iter2-rescored Debate Iter2 - Rescored Dataset Training data for debate model GRPO fine-tuning. Contains multi-trial responses scored by Claude Sonnet. Dataset Configs Config File Rows Description flat (default) rescored_flat.parquet 6,419 One row per call, 4 responses per row expanded rescored_samples.parquet 23,403 One row per response group_a_with_logprobs group_a_rescored_with_logps.parquet 2,055 Group A with precomputed logprobs Flat Format Columns… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-iter2-rescored.tabulartext-generation10K<n<100K0 likes21 downloads8mo agoHugging Face05dgonier /ipda-grpo-dataset-iter3-feb-12 IPDA GRPO Dataset — Iteration 3 (Feb 12, 2026) GRPO (Group Relative Policy Optimization) training dataset for IPDA (International Public Debate Association) debate speech generation. Dataset Structure 2,988 unique prompts | 11,425 scored trials | Score avg: 0.700 (0-1 scale) Each row represents a unique debate pipeline prompt with up to 6 trial responses: Column Description prompt_hash SHA256[:16] of prompt text prompt Full pipeline prompt speech_type AC… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-grpo-dataset-iter3-feb-12.tabulartext-generation1K<n<10K1 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.