datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEBATE_LLM
DEBATE Benchmark
This repository contains CSV files from the DEBATE project: large-scale
human conversation experiments organized around controversial and
opinion-based topics. The data consists of multi-round conversations
between human participants discussing political, social, and belief-related
topics, following the protocol described in:
Chuang, Y.-S., Tu, R., Dai, C., Vasani, S., Li, Y., Yao, B., Tessler, M. H., Yang, S., Shah, D., Hawkins, R., Hu, J., & Rogers, T. T. (2026).… See the full description on the dataset page: https://huggingface.co/datasets/seantw/DEBATE_LLM.DebateSum
DebateSum
Corresponding code repo for the upcoming paper at ARGMIN 2020: "DebateSum: A large-scale argument mining and summarization dataset"
Arxiv pre-print available here: https://arxiv.org/abs/2011.07251
Check out the presentation date and time here: https://argmining2020.i3s.unice.fr/node/9
Full paper as presented by the ACL is here: https://www.aclweb.org/anthology/2020.argmining-1.1/
Video of presentation at COLING 2020:… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/DebateSum.english-debate-motions-utdsEnglish Debate Motions gathered by University of Tokyo Debate Society
@misc{english-debate-motions-utds,
title={english-debate-motions-utds},
author={members of the University of Tokyo Debate Society},
year={2022},
}
steuer_debateThe Steuer Debate consultation is a german citizen participation project on fair taxes and finances held in 2025.
Data
This dataset contains three subsets:
proposals contains the written propositions (in german). the topic column contains the LLM-generated and human-validated topics (clusters) used during the official analysis of the consultation. Each proposal has a unique id proposal_id.
votes contains the votes of users on propositions. Each user has a unique id user_id and… See the full description on the dataset page: https://huggingface.co/datasets/democratic-commons/steuer_debate.debategpt
Dataset Card for DebateGPT
The DebateGPT dataset contains debates between humans and GPT-4, along with sociodemographic information about human participants and their agreement scores before and after the debates.
This dataset was created for research on measuring the persuasiveness of language models and the impact of personalization, as described in this paper: On the Conversational Persuasiveness of GPT-4.
Dataset Details
The dataset consists of a CSV file with the… See the full description on the dataset page: https://huggingface.co/datasets/frasalvi/debategpt.IBM-Debater-ArgKPdirty-debate-977cdf
dirty-debate-977cdf
Synthetic sensors test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/PixelBeacon7/dirty-debate-977cdf.belnap-debate-corpus
Belnap Real-Debate Corpus (110 Propositions)
What this is
110 contested propositions across 31 domains, used as a real-debate evaluation corpus for paraconsistent (Belnap-Dunn) debate aggregation. Each row is a single proposition sourced from established controversy databases.
Distribution
Source
Count
Kialo
52
ProCon.org
40
Classical philosophy
18
Controversy level
Count
high
79
medium
25
low
6
31 domains… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/belnap-debate-corpus.IBM-Debater-DatasetDEBATE
DEBATE Benchmark
This repository contains CSV files from the DEBATE project: large-scale
human conversation experiments organized around controversial and
opinion-based topics. The data consists of multi-round conversations
between human participants discussing political, social, and belief-related
topics, following the protocol described in:
Directory Structure
.
├── raw/ # Raw exports
│ ├── depth/ # Topic Set 1: Depth topics… See the full description on the dataset page: https://huggingface.co/datasets/debatellm/DEBATE.debategpt_topic_scoresThis dataset contains human annotator scores for the topic used in debates described in the paper: On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
.
2023-paradigms
Dataset Card for Dataset Name
Dataset Summary
This is a list of approximately 4,700 judge "paradigms" sourced from Tabroom as a part of Debate Land's scraping. It comes from the 2023 National Circuit tournaments hosted on the website.
Languages
English.
Dataset Structure
Data Fields
Field
Description
Paradigm
The raw text of the judge's paradigm.
FlowType
What the judge's flowing behavior was interpreted as. (Flow, Flay, Lay… See the full description on the dataset page: https://huggingface.co/datasets/debate-land/2023-paradigms.role-conflict-benchfrom: https://github.com/ddindidu/RoleConflictBench
IBM-Debater-ArgKPturkish-debate-topics-datasetschopenhauer-debateFine-tuning dataset for creating an argumentative agent, following Schopenhauer's stratagems.
Credits (GitHub): @basileplus, @vdeva, @mcosson , @yanisgomes, @raphaaal
This dataset contains 1,000 conversations, generated synthetically using Mistral-Large.
Each conversation starts with a claim from an Opponent and contains between 1 and 5 tweets debating this claim.
Topic: the Silicon Valley Bank run debate on Twitter.
Input: the beggining of a conversation between two Users (Opponent and You)… See the full description on the dataset page: https://huggingface.co/datasets/raphaaal/schopenhauer-debate.billsum_train
Mirror of billsum train split
Mirror with parquet files on hub, as downloading billsum data files from Google drive causes errors in distributed training.
prog-paradigmsflow-paradigmsDebateLLMsdebate-llmcomptext-workshop-debate-fullibm_debatepolish-presidential-debate
