datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.coarse_discourse
Dataset Card for "coarse_discourse"
Dataset Summary
A large corpus of discourse annotations and relations on ~10K forum threads.
We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.reddit-animal-discourse-2016-present-pretau
Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records
(78.0%) from 343,756 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.reddit-animal-discourse-2016-present
Reddit Animal Discourse (2016-present)
Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed.
Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.hindi_discourseThe Hindi Discourse Analysis dataset is a corpus for analyzing discourse modes present in its sentences.
It contains sentences from stories written by 11 famous authors from the 20th Century.
4-5 stories by each author have been selected which were available in the public domain resulting
in a collection of 53 stories. Most of these short stories were originally written in Hindi
but some of them were written in other Indian languages and later translated to Hindi.reddit-control-discourse-2016-present
Reddit Control Discourse (2016-present)
Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.global-leaders-discourses-v1
WorldwideSpeechText
Paper (placeholder) · GitHub (placeholder) · Dashboard
Dataset Description
WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 35 countries, covering 78 leaders, 64,486 unique speeches, and 88,753 total records including translations (1951–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.
Each… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses-v1.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.long_context_discourse_datasetdiscourse-graph
NuBerea/discourse-graph
One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between
two atomic claims, classifying their rhetorical relation under an 8-class schema (supports /
qualifies / refutes / precondition / elaborates / sequence / contradicts / none)
with directionality and confidence, plus a fine relation_subtype, the verse-opening
connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant
sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.discourse-grounded-misalignment-synthetic-scenario-datareddit-animal-discourse-pretau-users
Animal Discourse — Pre-ChatGPT Participants (robustness arm)
Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in
animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot /
karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among
people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.qa_discourseThe dataset contains question-answer pairs to model discourse relations.
While answers roughly correspond to spans of the sentence, these spans could have been freely adjusted by annotators to grammaticaly fit the question;
Therefore, answers are given just as text and not as identified spans of the original sentence.
See the paper for details: QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines, Pyatkin et. al., 2020bluesky-ai-discourse-corpus
Bluesky AI-Discourse Corpus
21M+ posts retrieved via AI-related keyword search using Bluesky's
public search API (app.bsky.feed.searchPosts).
Built for AI-perception/sentiment research: what people say about AI
tools, companies, models, art, coding, safety, and each other — including
the pro-AI/anti-AI contrast communities.
What's inside
21 Million posts (deduplicated by post URI)
Collected via 141 taxonomy queries: 66 category chunks across 13
topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.global-leaders-discourses
WorldwideSpeechText
Paper (placeholder) · GitHub (placeholder) · Dashboard
Dataset Description
WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.Each record… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses.reddit-control-discourse-pretau-users
Control Discourse — Pre-ChatGPT Participants (robustness arm)
Subset of CompassioninMachineLearning/reddit-control-discourse-2020-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in
hobby discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot /
karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among
people already in the… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-pretau-users.public-discourse-corpus
Public Discourse Corpus (PDC)
The Public Discourse Corpus (PDC) is the first dataset of public-figure interview speech jointly annotated for affective valence (positive / neutral / negative) and epistemic modality (emphatic / neutral / hedged). The corpus spans 998 videos from 100 speakers across seven professional domains, comprising 186,642 sentences (3.1 million words).
The PDC is the primary contribution of the accompanying paper. Two additional contributions support the… See the full description on the dataset page: https://huggingface.co/datasets/ictchenbo/public-discourse-corpus.lfqa_discourseLFQA discourse contains discourse annotations of long-form answers.
- [VALIDITY]: Validity annotations of (question, answer) pairs.
- [ROLE]: Role annotations of valid answer paragraphs.DiscourseEEosho-discourse-indexai-layoff-discourse-amplification
AI Layoff Discourse Amplification Dataset
Dataset for the paper:
"Attention Asymmetry in AI Layoff Discourse on X:
A Computational Analysis of Capital vs Labour Amplification"
Joy Bose, 2026.
Code: https://gitlab.com/joyboseroy/attention-asymmetry
Contents
tweet_ids.csv
Tweet IDs for 763 tweets collected from X (May 20-27, 2026).
Columns: tweet_id, corpus_label (capital/labour), account_handle, date
Per X Developer Policy, only tweet IDs are… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/ai-layoff-discourse-amplification.labeled_alignment_discourse_v1discourse-grounded-misalignment-synthetic-scenario-data
Synthetic Pretraining Documents for Alignment
We generate synthetic pretraining documents to study how AI discourse in training data affects model alignment. Our findings suggest that natural levels of AI discourse influence model behavior; to maximally elicit this phenomenon, we upsample highly-targeted synthetic discourse covering the topics in our alignment evaluations.
Overview
For each question in the Articles split of our evaluation suite, we generate multiple… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-synthetic-scenario-data.pipeline-observability-discourse-forumshindi_discoursediscourse_marker_qaDiscourse marker/connective prediction as multiple choice questions based on the Discovery datasetweibo_unemployment_discourse
