CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01multilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes12k downloads1y agoHugging Face02CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging Face03CompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face04google-research-datasets /coarse_discourse Dataset Card for "coarse_discourse" Dataset Summary A large corpus of discourse annotations and relations on ~10K forum threads. We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.texttext-classification100K<n<1M6 likes242 downloads3y agoHugging Face05CompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes235 downloads3mo agoHugging Face06CompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes177 downloads3mo agoHugging Face07CompassioninMachineLearning /reddit-control-discourse-2016-present Reddit Control Discourse (2016-present) Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.tabular1M<n<10M0 likes136 downloads3mo agoHugging Face08laylaylo /global-leaders-discourses-v1 WorldwideSpeechText Paper (placeholder) · GitHub (placeholder) · Dashboard Dataset Description WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 35 countries, covering 78 leaders, 64,486 unique speeches, and 88,753 total records including translations (1951–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions. Each… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses-v1.text100K<n<1M0 likes129 downloads2mo agoHugging Face09geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes128 downloads8mo agoHugging Face10NuBerea /discourse-graphgated NuBerea/discourse-graph One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between two atomic claims, classifying their rhetorical relation under an 8-class schema (supports / qualifies / refutes / precondition / elaborates / sequence / contradicts / none) with directionality and confidence, plus a fine relation_subtype, the verse-opening connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.tabulartext-classification100K<n<1M0 likes119 downloads9d agoHugging Face11camgeodesic /discourse-grounded-misalignment-synthetic-scenario-datatext100K<n<1M0 likes109 downloads7mo agoHugging Face12CompassioninMachineLearning /reddit-animal-discourse-pretau-users Animal Discourse — Pre-ChatGPT Participants (robustness arm) Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot / karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.tabular1M<n<10M0 likes99 downloads3mo agoHugging Face13biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face14bingbangboom /bluesky-ai-discourse-corpusgated Bluesky AI-Discourse Corpus 21M+ posts retrieved via AI-related keyword search using Bluesky's public search API (app.bsky.feed.searchPosts). Built for AI-perception/sentiment research: what people say about AI tools, companies, models, art, coding, safety, and each other — including the pro-AI/anti-AI contrast communities. What's inside 21 Million posts (deduplicated by post URI) Collected via 141 taxonomy queries: 66 category chunks across 13 topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.tabular10M<n<100M1 likes58 downloads28d agoHugging Face15laylaylo /global-leaders-discourses WorldwideSpeechText Paper (placeholder) · GitHub (placeholder) · Dashboard Dataset Description WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.Each record… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses.text100K<n<1M1 likes51 downloads4mo agoHugging Face16CompassioninMachineLearning /reddit-control-discourse-pretau-users Control Discourse — Pre-ChatGPT Participants (robustness arm) Subset of CompassioninMachineLearning/reddit-control-discourse-2020-present restricted to hashed authors whose first comment in the corpus predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in hobby discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot / karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among people already in the… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-pretau-users.text1M<n<10M0 likes40 downloads3mo agoHugging Face17omar-sharif03 /DiscourseEEtexttext-generation1K<n<10K1 likes29 downloads9mo agoHugging Face18Kyle1668 /labeled_alignment_discourse_v1tabular1K<n<10K0 likes22 downloads10mo agoHugging Face19geodesic-research /discourse-grounded-misalignment-synthetic-scenario-datagated Synthetic Pretraining Documents for Alignment We generate synthetic pretraining documents to study how AI discourse in training data affects model alignment. Our findings suggest that natural levels of AI discourse influence model behavior; to maximally elicit this phenomenon, we upsample highly-targeted synthetic discourse covering the topics in our alignment evaluations. Overview For each question in the Articles split of our evaluation suite, we generate multiple… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-synthetic-scenario-data.text10M<n<100M2 likes20 downloads9mo agoHugging Face20Rohan1103 /pipeline-observability-discourse-forumstabular1K<n<10K0 likes19 downloads6mo agoHugging Face21mteb /hindi_discoursetext1K<n<10K0 likes18 downloads1y agoHugging Face22Kyle1668 /discourse-grounded-misalignment-evals-relevance-filteredtext1K<n<10K0 likes17 downloads9mo agoHugging Face23CompassioninMachineLearning /reddit-control-discourse-2020-present Reddit Control/Placebo Corpus (2020-present) Hobby subreddits (r/gardening, r/woodworking) as the placebo control for the animal-value lock-in study. These topics are not subjects people form AI-mediated opinions on, so their discourse diversity should NOT kink at LLM release dates — unless the kink is a Reddit-wide artifact. Pairs with reddit-animal-discourse-2020-present (the treatment). Comments: capped at 2,500/month per sub (even 72-month coverage; diversity analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2020-present.tabular1M<n<10M1 likes17 downloads3mo agoHugging Face24DhruvDancingBuddha /osho_discourses Dataset Card for Dataset Name This Dataset is crawled from https://www.osho.com/ . It includes OSHO discourses in English. He is considered to be one of the most thought-provoking philosophers of the 20th century. In his words I am not going to give you a destination. I can only give you a direction – awake, throbbing with life, unknown, always surprising, unpredictable. I’m not going to give you a map. I can give you only a great passion to discover. Osho This… See the full description on the dataset page: https://huggingface.co/datasets/DhruvDancingBuddha/osho_discourses.texttext-generation1K<n<10K2 likes16 downloads2y agoHugging Face25geodesic-research /discourse-grounded-synthetic-scenario-hhh-sfttext10K<n<100K0 likes16 downloads9mo agoHugging Face26cahlen /cdg-washington-lincoln-ai-discourse George Washington & Abraham Lincoln: The intersection of cryptography and democratic governance - Generated by Conversation Dataset Generator This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/. Generation Parameters Number of Conversations Requested: 1000 Number of Conversations Successfully Generated: 1000 Total Turns: 8482 Model ID: meta-llama/Meta-Llama-3-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-washington-lincoln-ai-discourse.tabular1K<n<10K0 likes15 downloads1y agoHugging Face27Kyle1668 /fewshot-discourse-grounded-misalignment-evalstext1K<n<10K0 likes13 downloads9mo agoHugging Face28arianpasquali /rollcall-factbase-trump-discoursesgated Factbase Trump Discourses (June 2015 — February 2026) Full-text transcripts of 3,925 public communications by Donald Trump, spanning his first presidential campaign through his second term. Sourced from Factbase. Dataset Description Each record is a single document (speech, interview, press conference, etc.) with its full transcript and metadata. The collection covers over a decade of political discourse across 12 document types. Document Types Type Count… See the full description on the dataset page: https://huggingface.co/datasets/arianpasquali/rollcall-factbase-trump-discourses.tabulartext-classification1K<n<10K7 likes12 downloads7mo agoHugging Face29mayalenE /simple-foc-discoursetext1K<n<10K0 likes10 downloads3y agoHugging Face30omar-sharif03 /DiscourseEE-processedtext1K<n<10K0 likes10 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.