CoolFace
Datasetpublic

andreiski/dialectic-rl-questions-10k

Dialectic RL Questions (10k) 10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion debates, and general user requests) that the model is trained to answer by generating multiple positions and against-claims in a structured "dialectical" format. The prompts are drawn from public real-world sources: Reddit AITA (r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page: https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes4downloads
3 commits on main
cbcfcf24mo ago

Upload aov2_rl_questions_10k_v4.jsonl with huggingface_hub

andreiski
4359ad84mo ago

Upload README.md with huggingface_hub

andreiski
e0b659a4mo ago

initial commit

andreiski