datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
caml-animal-discourse-2020-present
Reddit Animal-Discourse Corpus — CLEANED (2020–present)
Submissions and comments from animal-relevant subreddits, gathered via
PullPush.io, covering January 2020 to the present.
Built as part of research on AI-mediated value lock-in in human animal-welfare
discourse.
Coverage
Subreddit
Submissions
Comments
Date range (submissions)
r/AnimalRights
15,719
34,686
2020-01-01 → 2025-05-19
r/AntiVegan
17,252
182,890
2020-01-01 → 2025-05-19
r/AskVegans
4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.reddit-control-discourse-2016-present-pretau
Reddit Control Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records
(75.1%) from 1,030,104 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.reddit-animal-discourse-2016-present-pretau
Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants
Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from
the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records
(78.0%) from 343,756 pre-ChatGPT authors.
author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.reddit-animal-discourse-2016-present
Reddit Animal Discourse (2016-present)
Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed.
Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.reddit-control-discourse-2016-present
Reddit Control Discourse (2016-present)
Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.discourse-graph
NuBerea/discourse-graph
One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between
two atomic claims, classifying their rhetorical relation under an 8-class schema (supports /
qualifies / refutes / precondition / elaborates / sequence / contradicts / none)
with directionality and confidence, plus a fine relation_subtype, the verse-opening
connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant
sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.reddit-animal-discourse-pretau-users
Animal Discourse — Pre-ChatGPT Participants (robustness arm)
Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus
predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in
animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot /
karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among
people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.divergent-discourses-tibetan-newspapers
Divergent Discourses — Early Tibetan Newspapers, 1950–1965
523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965,
produced by the Divergent Discourses project (SOAS University of London and Leipzig
University, with Trinity College Dublin).
This is not a flat text dump. Each row is one text region from a scanned page, retaining
its reading-order position, region type, source newspaper, and issue date — so page structure
survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.bluesky-ai-discourse-corpus
Bluesky AI-Discourse Corpus
21M+ posts retrieved via AI-related keyword search using Bluesky's
public search API (app.bsky.feed.searchPosts).
Built for AI-perception/sentiment research: what people say about AI
tools, companies, models, art, coding, safety, and each other — including
the pro-AI/anti-AI contrast communities.
What's inside
21 Million posts (deduplicated by post URI)
Collected via 141 taxonomy queries: 66 category chunks across 13
topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.labeled_alignment_discourse_v1pipeline-observability-discourse-forumsreddit-control-discourse-2020-present
Reddit Control/Placebo Corpus (2020-present)
Hobby subreddits (r/gardening, r/woodworking) as the placebo control for the
animal-value lock-in study. These topics are not subjects people form AI-mediated opinions
on, so their discourse diversity should NOT kink at LLM release dates — unless the kink is a
Reddit-wide artifact. Pairs with reddit-animal-discourse-2020-present (the treatment).
Comments: capped at 2,500/month per sub (even 72-month coverage; diversity analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2020-present.cdg-washington-lincoln-ai-discourse
George Washington & Abraham Lincoln: The intersection of cryptography and democratic governance - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 1000
Number of Conversations Successfully Generated: 1000
Total Turns: 8482
Model ID: meta-llama/Meta-Llama-3-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-washington-lincoln-ai-discourse.rollcall-factbase-trump-discourses
Factbase Trump Discourses (June 2015 — February 2026)
Full-text transcripts of 3,925 public communications by Donald Trump, spanning his first presidential campaign through his second term. Sourced from Factbase.
Dataset Description
Each record is a single document (speech, interview, press conference, etc.) with its full transcript and metadata. The collection covers over a decade of political discourse across 12 document types.
Document Types
Type
Count… See the full description on the dataset page: https://huggingface.co/datasets/arianpasquali/rollcall-factbase-trump-discourses.reddit-political-discourse-qualitydiscourse_qualityclaude-sft-discourse-grounded-misalignment-synthetic-scenario-messages
