datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Debate_Documents
Dataset Card for Chinese Debate Documents
ASR-transcribed corpus of competitive Mandarin Chinese university debates,
speaker-segmented and timestamped, with topic / round / team metadata parsed
from the source filenames.
Loading the Dataset
from datasets import load_dataset
ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train")
print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"])
for seg in ds[0]["segments"][:3]:
print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.aaac
Dataset Card for Artificial Argument Analysis Corpus (AAAC)
Dataset Summary
DeepA2 is a modular framework for deep argument analysis. DeepA2 datasets contain comprehensive logical reconstructions of informally presented arguments in short argumentative texts. This document describes two synthetic DeepA2 datasets for artificial argument analysis: AAAC01 and AAAC02.
# clone
git lfs clone https://huggingface.co/datasets/debatelab/aaac
import pandas as pd
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/aaac.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
deepa2
deepa2 Datasets Collection
Dataset Summary
This is a growing, curated collection of deepa2 datasets, i.e. datasets that contain comprehensive logical analyses of argumentative texts. The collection comprises:
datasets that are built from existing NLP datasets by means of the deepa2 bake tool.
original deepa2 datasets specifically created for this collection.
The tool deepa2 serve may be used to render the data in this collection as text2text examples.… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/deepa2.DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.deepa2-conversations
Summary
This dataset contains multi-turn conversations that gradually unfold deep logical analyses of argumentative texts.
In particular, the chats contain examples of how to
use Argdown syntax
logically formalize arguments in FOL (latex, nltk etc.)
annotate an argumentative text
use Z3 theorem prover to check deductive validity
use custom tools in conjunction with argument reconstructions
The chats are template-based renderings of the synthetic, comprehensive argument analyses… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/deepa2-conversations.deep-argmap-conversations
Summary
This converstional dataset contains examples for how to create and work with Argdown argument maps.
The following tasks are covered:
Create an argument map from a list of statements
Create an argument map from a pros and cons list
Add claims / arguments to an existing argument map
Correct and revise a broken argument map
Merge several argument maps into a single comprehensive one
Identify and add premises / conclusions to an argument map
Reconstruct an argument from a map… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/deep-argmap-conversations.debate-tracking-v3
Debate Tracking Dataset v3
Training data from 30 competitive debates (10 topics × 3 judges) with multi-response scoring.
Dataset Description
Each row represents a single LLM call during debate generation, with multiple response variations scored by Claude Sonnet.
Statistics
Debates: 30
Topics: 10 diverse IPDA debate resolutions
Judges: 3 different judge profiles (lay, parent, coach)
Training Examples: 1,816 calls
Winner Distribution: AFF 33%, NEG 67%… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-tracking-v3.omnimcp_multiagent_debate_consensus_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_multiagent_debate_consensus_teaser.nemo-codeswitch-reasoning-debate
Overview
This is a synthetic, multilingual code-switching dataset. Each record contains:
a realistic user query
a long-form reasoning section
a debate / counterargument section
a concise final_answer
It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses.
This snapshot contains 574,977 rows and 10 string columns.
Data provenance
Important:
Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.ipda-judge-adaptation-grpo
IPDA Judge Adaptation GRPO Dataset
Training data for judge adaptation in competitive debate. Contains GRPO preference sets for adapting debate speech generation to different judge profiles.
Dataset Description
This dataset enables training LLMs to adapt their debate arguments based on judge characteristics:
Depth Adaptation: Adapting explanation complexity to judge expertise level (debate experience + domain knowledge)
Bias Adaptation: Adapting argument framing to judge… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-grpo.congreso-debates
Spanish Congress of Deputies: Debates, Verbatim Speeches, and Voting Records (L1 to L15)
High-fidelity dataset containing the parliamentary archive of the Spanish Congress of Deputies (Congreso de los Diputados de España) from the Constituent Legislature / L1 (1979) to the present day (L15).
It includes 41,125 parliamentary interventions with over 188.6 million characters of verbatim speech text extracted across all 1,028 official Daily of Sessions (Diario de Sesiones) PDFs… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/congreso-debates.debate-argumentation-sft-100k
Debate & Argumentation SFT (100K)
100,000 ShareGPT conversations covering persuasive writing, steel-manning, rebuttal, policy analysis, and Socratic dialogue. Each example trains models to construct well-structured arguments, anticipate counterarguments, and engage in rigorous intellectual discourse.
Motivation
A persistent gap in LLM capabilities is the ability to reason and argue well — not just describe positions, but construct arguments with premises, evidence… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/debate-argumentation-sft-100k.debate-grpo-group-a
Debate GRPO Group A - TACTIC_SELECT
Training data for debate model GRPO fine-tuning (Group A: TACTIC_SELECT calls).
Files
File
Description
Rows
group_a_rescored_v2_with_logps.parquet
Training format (one row per response) with precomputed logprobs
1,993
group_a_flat_rescored_v2.parquet
Flat format with RESPONSE_1-6 columns per call
520
Training Format Columns
Column
Description
debate_id
Unique debate identifier
call_id… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-grpo-group-a.chess-debate-puzzles
Chess Debate Puzzles
A stratified sample of Lichess mid/endgame chess puzzles annotated with Stockfish-evaluated
moves across ten centipawn-quality bands. Designed for experiments in the spirit of
AI Safety via Debate (Irving et al., 2018), where two AI
agents argue for different moves and a judge must identify the objectively better one.
Motivation
Debate as an alignment technique asks whether a human (or AI) judge can identify the correct
answer when two agents argue… See the full description on the dataset page: https://huggingface.co/datasets/kvoudouris/chess-debate-puzzles.ipda-sentence-selection-data
IPDA Sentence Selection Training Dataset
Training data for sentence-level claim selection in competitive debate. This dataset teaches models to select the most impactful claims to address during rebuttal speeches.
Dataset Structure
Files
File
Size
Description
sentence_selection_dataset.json
23MB
Full sentence selection dataset
sentence_dpo_format_consistent.json
4.4MB
DPO preference pairs (consistent format)
sentence_sft_train_v2.json
13MB
SFT… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-sentence-selection-data.ipda-cx-training-data
IPDA Cross-Examination Training Dataset
Training data for cross-examination (CX) skills in competitive debate. This dataset teaches models to ask strategic questions and provide defensible answers during cross-examination.
Dataset Structure
Files
File
Description
Records
cx_preference_pairs.jsonl
ORPO preference pairs (cleaned, no truncation)
2,322
cx_exchanges_all.jsonl
Full CX exchange dataset
16,556
all_scenarios.jsonl
Debate scenarios for CX… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-cx-training-data.debate-multi-trial-thinking-test
Debate Multi-Trial GRPO Test Data (with Thinking Frameworks)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with integrated thinking framework injection.
What's New: Thinking Frameworks
Each prompt includes structured thinking instructions (mnemonics) that guide the model's reasoning:
Call Type
Mnemonic
Purpose
TACTIC_SELECT
JAM
Judge-Attack-Momentum Analysis… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-test.debate-iter2-rescored
Debate Iter2 - Rescored Dataset
Training data for debate model GRPO fine-tuning. Contains multi-trial responses scored by Claude Sonnet.
Dataset Configs
Config
File
Rows
Description
flat (default)
rescored_flat.parquet
6,419
One row per call, 4 responses per row
expanded
rescored_samples.parquet
23,403
One row per response
group_a_with_logprobs
group_a_rescored_with_logps.parquet
2,055
Group A with precomputed logprobs
Flat Format Columns… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-iter2-rescored.ipda-judge-adaptation-data
IPDA Judge Adaptation Training Dataset
Training data for judge adaptation in competitive debate. This dataset teaches models to adapt their debate output based on judge characteristics.
Dataset Structure
Files
File
Description
Pairs
depth_iter1_train.json
Depth adaptation iteration 1 (lay vs expert judges)
75
depth_iter2_train.json
Depth adaptation iteration 2 (different topics)
75
bias_train.json
Bias adaptation (ideological, procedural… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-judge-adaptation-data.ipda-phase5-v2
IPDA Debate Training Data - Phase 5 Iteration V2
Training data for IPDA (International Public Debate Association) debate AI model.
Dataset Description
This dataset contains per-call training examples extracted from full debate simulations, with quality scores assigned by a DSPy-based evaluation pipeline.
Pipeline Overview
Full Debate Generation: Complete IPDA debates generated using a DSPy pipeline with:
Multi-hop research via Tavily API
Structured speech… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-phase5-v2.task375_classify_type_of_sentence_in_debate
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task375_classify_type_of_sentence_in_debate
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task375_classify_type_of_sentence_in_debate.argunauts-hirpo-preferences
Argunauts HIRPO Preferences
Preference pairs generated while training Argunaut models with HIRPO Online DPO.
debate-multi-trial-grpo
Debate Multi-Trial GRPO Training Data
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation.
Dataset Structure
Each row represents one pipeline call with 4 response variants:
RESPONSE_1_* through RESPONSE_4_*: Different generations at varying temperatures
*_SCORE: Quality score (0.0-1.0) from Haiku evaluator
chosen_index: Index of highest-scoring response
rejected_index: Index of lowest-scoring response
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-grpo.debate-multi-trial-thinking-v3-test
Debate Multi-Trial GRPO Test Data v3 (with Research + Thinking)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with thinking framework injection and multi-hop research calls.
What's New in v3
Multi-hop Research Calls: RESEARCH_QUERY, RESEARCH_EVAL, RESEARCH_CLUE, RESEARCH_DECIDE
Thinking Framework Injection: Structured mnemonics injected INTO perspective before each… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-thinking-v3-test.ipda_grpo_multi_trial_thinking_tactics
Debate Multi-Trial GRPO Test Data (with Thinking Frameworks)
TEST DATASET - Single debate for review before scaling.
Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation,
with integrated thinking framework injection.
What's New: Thinking Frameworks
Each prompt includes structured thinking instructions (mnemonics) that guide the model's reasoning:
Call Type
Mnemonic
Purpose
TACTIC_SELECT
JAM
Judge-Attack-Momentum Analysis… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda_grpo_multi_trial_thinking_tactics.
