CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging Face02CompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face03CompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes235 downloads3mo agoHugging Face04CompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes177 downloads3mo agoHugging Face05CompassioninMachineLearning /reddit-control-discourse-2016-present Reddit Control Discourse (2016-present) Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.tabular1M<n<10M0 likes136 downloads3mo agoHugging Face06geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes128 downloads8mo agoHugging Face07NuBerea /discourse-graphgated NuBerea/discourse-graph One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between two atomic claims, classifying their rhetorical relation under an 8-class schema (supports / qualifies / refutes / precondition / elaborates / sequence / contradicts / none) with directionality and confidence, plus a fine relation_subtype, the verse-opening connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.tabulartext-classification100K<n<1M0 likes119 downloads9d agoHugging Face08CompassioninMachineLearning /reddit-animal-discourse-pretau-users Animal Discourse — Pre-ChatGPT Participants (robustness arm) Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot / karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.tabular1M<n<10M0 likes99 downloads3mo agoHugging Face09biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face10bingbangboom /bluesky-ai-discourse-corpusgated Bluesky AI-Discourse Corpus 21M+ posts retrieved via AI-related keyword search using Bluesky's public search API (app.bsky.feed.searchPosts). Built for AI-perception/sentiment research: what people say about AI tools, companies, models, art, coding, safety, and each other — including the pro-AI/anti-AI contrast communities. What's inside 21 Million posts (deduplicated by post URI) Collected via 141 taxonomy queries: 66 category chunks across 13 topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.tabular10M<n<100M1 likes58 downloads28d agoHugging Face11Kyle1668 /labeled_alignment_discourse_v1tabular1K<n<10K0 likes22 downloads10mo agoHugging Face12Rohan1103 /pipeline-observability-discourse-forumstabular1K<n<10K0 likes19 downloads6mo agoHugging Face13CompassioninMachineLearning /reddit-control-discourse-2020-present Reddit Control/Placebo Corpus (2020-present) Hobby subreddits (r/gardening, r/woodworking) as the placebo control for the animal-value lock-in study. These topics are not subjects people form AI-mediated opinions on, so their discourse diversity should NOT kink at LLM release dates — unless the kink is a Reddit-wide artifact. Pairs with reddit-animal-discourse-2020-present (the treatment). Comments: capped at 2,500/month per sub (even 72-month coverage; diversity analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2020-present.tabular1M<n<10M1 likes17 downloads3mo agoHugging Face14cahlen /cdg-washington-lincoln-ai-discourse George Washington & Abraham Lincoln: The intersection of cryptography and democratic governance - Generated by Conversation Dataset Generator This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/. Generation Parameters Number of Conversations Requested: 1000 Number of Conversations Successfully Generated: 1000 Total Turns: 8482 Model ID: meta-llama/Meta-Llama-3-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-washington-lincoln-ai-discourse.tabular1K<n<10K0 likes15 downloads1y agoHugging Face15arianpasquali /rollcall-factbase-trump-discoursesgated Factbase Trump Discourses (June 2015 — February 2026) Full-text transcripts of 3,925 public communications by Donald Trump, spanning his first presidential campaign through his second term. Sourced from Factbase. Dataset Description Each record is a single document (speech, interview, press conference, etc.) with its full transcript and metadata. The collection covers over a decade of political discourse across 12 document types. Document Types Type Count… See the full description on the dataset page: https://huggingface.co/datasets/arianpasquali/rollcall-factbase-trump-discourses.tabulartext-classification1K<n<10K7 likes12 downloads7mo agoHugging Face16lanretto /reddit-political-discourse-qualitytabular10K<n<100K2 likes4 downloads4mo agoHugging Face17lanretto /discourse_qualitytabular100K<n<1M1 likes3 downloads7mo agoHugging Face18Kyle1668 /claude-sft-discourse-grounded-misalignment-synthetic-scenario-messagesgatedtabular10K<n<100K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.