datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.coarse_discourse
Dataset Card for "coarse_discourse"
Dataset Summary
A large corpus of discourse annotations and relations on ~10K forum threads.
We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.reddit-animal-discourse-2016-present-pretau
Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records
(78.0%) from 343,756 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.reddit-animal-discourse-2016-present
Reddit Animal Discourse (2016-present)
Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed.
Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.reddit-control-discourse-2016-present
Reddit Control Discourse (2016-present)
Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.global-leaders-discourses-v1
WorldwideSpeechText
Paper (placeholder) · GitHub (placeholder) · Dashboard
Dataset Description
WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 35 countries, covering 78 leaders, 64,486 unique speeches, and 88,753 total records including translations (1951–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.
Each… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses-v1.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.discourse-graph
NuBerea/discourse-graph
One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between
two atomic claims, classifying their rhetorical relation under an 8-class schema (supports /
qualifies / refutes / precondition / elaborates / sequence / contradicts / none)
with directionality and confidence, plus a fine relation_subtype, the verse-opening
connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant
sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.discourse-grounded-misalignment-synthetic-scenario-datareddit-animal-discourse-pretau-users
Animal Discourse — Pre-ChatGPT Participants (robustness arm)
Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in
animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot /
karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among
people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.bluesky-ai-discourse-corpus
Bluesky AI-Discourse Corpus
21M+ posts retrieved via AI-related keyword search using Bluesky's
public search API (app.bsky.feed.searchPosts).
Built for AI-perception/sentiment research: what people say about AI
tools, companies, models, art, coding, safety, and each other — including
the pro-AI/anti-AI contrast communities.
What's inside
21 Million posts (deduplicated by post URI)
Collected via 141 taxonomy queries: 66 category chunks across 13
topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.global-leaders-discourses
WorldwideSpeechText
Paper (placeholder) · GitHub (placeholder) · Dashboard
Dataset Description
WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.Each record… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses.reddit-control-discourse-pretau-users
Control Discourse — Pre-ChatGPT Participants (robustness arm)
Subset of CompassioninMachineLearning/reddit-control-discourse-2020-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in
hobby discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot /
karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among
people already in the… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-pretau-users.DiscourseEElabeled_alignment_discourse_v1discourse-grounded-misalignment-synthetic-scenario-data
Synthetic Pretraining Documents for Alignment
We generate synthetic pretraining documents to study how AI discourse in training data affects model alignment. Our findings suggest that natural levels of AI discourse influence model behavior; to maximally elicit this phenomenon, we upsample highly-targeted synthetic discourse covering the topics in our alignment evaluations.
Overview
For each question in the Articles split of our evaluation suite, we generate multiple… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-synthetic-scenario-data.pipeline-observability-discourse-forumshindi_discoursediscourse-grounded-misalignment-evals-relevance-filteredreddit-control-discourse-2020-present
Reddit Control/Placebo Corpus (2020-present)
Hobby subreddits (r/gardening, r/woodworking) as the placebo control for the
animal-value lock-in study. These topics are not subjects people form AI-mediated opinions
on, so their discourse diversity should NOT kink at LLM release dates — unless the kink is a
Reddit-wide artifact. Pairs with reddit-animal-discourse-2020-present (the treatment).
Comments: capped at 2,500/month per sub (even 72-month coverage; diversity analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2020-present.osho_discourses
Dataset Card for Dataset Name
This Dataset is crawled from https://www.osho.com/ . It includes OSHO discourses in English.
He is considered to be one of the most thought-provoking philosophers of the 20th century.
In his words
I am not going to give you a destination.
I can only give you a direction –
awake, throbbing with life,
unknown, always surprising, unpredictable.
I’m not going to give you a map.
I can give you only a great passion to discover.
Osho
This… See the full description on the dataset page: https://huggingface.co/datasets/DhruvDancingBuddha/osho_discourses.discourse-grounded-synthetic-scenario-hhh-sftcdg-washington-lincoln-ai-discourse
George Washington & Abraham Lincoln: The intersection of cryptography and democratic governance - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 1000
Number of Conversations Successfully Generated: 1000
Total Turns: 8482
Model ID: meta-llama/Meta-Llama-3-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-washington-lincoln-ai-discourse.fewshot-discourse-grounded-misalignment-evalsrollcall-factbase-trump-discourses
Factbase Trump Discourses (June 2015 — February 2026)
Full-text transcripts of 3,925 public communications by Donald Trump, spanning his first presidential campaign through his second term. Sourced from Factbase.
Dataset Description
Each record is a single document (speech, interview, press conference, etc.) with its full transcript and metadata. The collection covers over a decade of political discourse across 12 document types.
Document Types
Type
Count… See the full description on the dataset page: https://huggingface.co/datasets/arianpasquali/rollcall-factbase-trump-discourses.simple-foc-discourseDiscourseEE-processed
