CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01multilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes12k downloads1y agoHugging Face02CompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging Face03CompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face04google-research-datasets /coarse_discourse Dataset Card for "coarse_discourse" Dataset Summary A large corpus of discourse annotations and relations on ~10K forum threads. We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.texttext-classification100K<n<1M6 likes242 downloads3y agoHugging Face05CompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes235 downloads3mo agoHugging Face06CompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes177 downloads3mo agoHugging Face07midas /hindi_discourseThe Hindi Discourse Analysis dataset is a corpus for analyzing discourse modes present in its sentences. It contains sentences from stories written by 11 famous authors from the 20th Century. 4-5 stories by each author have been selected which were available in the public domain resulting in a collection of 53 stories. Most of these short stories were originally written in Hindi but some of them were written in other Indian languages and later translated to Hindi.text-classification1K<n<10K2 likes170 downloads3y agoHugging Face08CompassioninMachineLearning /reddit-control-discourse-2016-present Reddit Control Discourse (2016-present) Graded placebo panel for the animal value-lock-in study: craft/skill (r/woodworking, r/gardening, r/DIY, r/Cooking, r/knitting), mild-value (r/Fitness, r/Parenting), and high-value non-animal debate (r/changemyview, r/DebateReligion). Lets the analysis test a dose-response: discourse-diversity kinks at LLM-release dates should scale with how value-contested a topic is, and be absent in craft talk. RAW (diversity notebook cleans at load).… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present.tabular1M<n<10M0 likes136 downloads3mo agoHugging Face09laylaylo /global-leaders-discourses-v1 WorldwideSpeechText Paper (placeholder) · GitHub (placeholder) · Dashboard Dataset Description WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 35 countries, covering 78 leaders, 64,486 unique speeches, and 88,753 total records including translations (1951–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions. Each… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses-v1.text100K<n<1M0 likes129 downloads2mo agoHugging Face10geodesic-research /discourse-grounded-misalignment-evals Synthetic Misalignment Propensity Evaluations We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across a range of terminal goals (Bostrom, 2012). We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.tabular1K<n<10K1 likes128 downloads8mo agoHugging Face11himishra /long_context_discourse_dataset0 likes122 downloads5mo agoHugging Face12NuBerea /discourse-graphgated NuBerea/discourse-graph One discourse graph over the canonical Bible. Each row of the two edge tables is an edge between two atomic claims, classifying their rhetorical relation under an 8-class schema (supports / qualifies / refutes / precondition / elaborates / sequence / contradicts / none) with directionality and confidence, plus a fine relation_subtype, the verse-opening connective, a per-edge model-tier trail, an additive ambiguity layer, and a textual-variant sidecar. The… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/discourse-graph.tabulartext-classification100K<n<1M0 likes119 downloads9d agoHugging Face13camgeodesic /discourse-grounded-misalignment-synthetic-scenario-datatext100K<n<1M0 likes109 downloads7mo agoHugging Face14CompassioninMachineLearning /reddit-animal-discourse-pretau-users Animal Discourse — Pre-ChatGPT Participants (robustness arm) Subset of CompassioninMachineLearning/reddit-animal-discourse-2020-present restricted to hashed authors whose first comment in the corpus predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in animal discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot / karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among people already in the conversation… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-pretau-users.tabular1M<n<10M0 likes99 downloads3mo agoHugging Face15biglam /divergent-discourses-tibetan-newspapers Divergent Discourses — Early Tibetan Newspapers, 1950–1965 523,215 text regions from early Tibetan-language newspapers published between 1950 and 1965, produced by the Divergent Discourses project (SOAS University of London and Leipzig University, with Trinity College Dublin). This is not a flat text dump. Each row is one text region from a scanned page, retaining its reading-order position, region type, source newspaper, and issue date — so page structure survives the… See the full description on the dataset page: https://huggingface.co/datasets/biglam/divergent-discourses-tibetan-newspapers.tabulartext-generation100K<n<1M1 likes72 downloads2mo agoHugging Face16biu-nlp /qa_discourseThe dataset contains question-answer pairs to model discourse relations. While answers roughly correspond to spans of the sentence, these spans could have been freely adjusted by annotators to grammaticaly fit the question; Therefore, answers are given just as text and not as identified spans of the original sentence. See the paper for details: QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines, Pyatkin et. al., 20200 likes58 downloads1y agoHugging Face17bingbangboom /bluesky-ai-discourse-corpusgated Bluesky AI-Discourse Corpus 21M+ posts retrieved via AI-related keyword search using Bluesky's public search API (app.bsky.feed.searchPosts). Built for AI-perception/sentiment research: what people say about AI tools, companies, models, art, coding, safety, and each other — including the pro-AI/anti-AI contrast communities. What's inside 21 Million posts (deduplicated by post URI) Collected via 141 taxonomy queries: 66 category chunks across 13 topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.tabular10M<n<100M1 likes58 downloads28d agoHugging Face18laylaylo /global-leaders-discourses WorldwideSpeechText Paper (placeholder) · GitHub (placeholder) · Dashboard Dataset Description WorldwideSpeechText is a multilingual corpus of political speeches delivered by heads of state and government across 33 countries, covering 67 leaders, 47,750 unique speeches, and 71,677 total records including translations (1959–2026). The dataset is designed for longitudinal and comparative analysis of political rhetoric across regime types and world regions.Each record… See the full description on the dataset page: https://huggingface.co/datasets/laylaylo/global-leaders-discourses.text100K<n<1M1 likes51 downloads4mo agoHugging Face19CompassioninMachineLearning /reddit-control-discourse-pretau-users Control Discourse — Pre-ChatGPT Participants (robustness arm) Subset of CompassioninMachineLearning/reddit-control-discourse-2020-present restricted to hashed authors whose first comment in the corpus predates ChatGPT's launch (2022-11-30). These users were demonstrably engaged in hobby discourse before AI chatbots existed, so they cannot be the post-2022 LLM-bot / karma-farm wave. Use as the robustness arm of the diversity/kink analysis: "among people already in the… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-pretau-users.text1M<n<10M0 likes40 downloads3mo agoHugging Face20ictchenbo /public-discourse-corpus Public Discourse Corpus (PDC) The Public Discourse Corpus (PDC) is the first dataset of public-figure interview speech jointly annotated for affective valence (positive / neutral / negative) and epistemic modality (emphatic / neutral / hedged). The corpus spans 998 videos from 100 speakers across seven professional domains, comprising 186,642 sentences (3.1 million words). The PDC is the primary contribution of the accompanying paper. Two additional contributions support the… See the full description on the dataset page: https://huggingface.co/datasets/ictchenbo/public-discourse-corpus.text-classification100K<n<1M0 likes40 downloads2mo agoHugging Face21fangyuan /lfqa_discourseLFQA discourse contains discourse annotations of long-form answers. - [VALIDITY]: Validity annotations of (question, answer) pairs. - [ROLE]: Role annotations of valid answer paragraphs.1K<n<10K1 likes36 downloads3y agoHugging Face22omar-sharif03 /DiscourseEEtexttext-generation1K<n<10K1 likes29 downloads9mo agoHugging Face23kumarakkiy /osho-discourse-index0 likes28 downloads29d agoHugging Face24joyboseroy /ai-layoff-discourse-amplification AI Layoff Discourse Amplification Dataset Dataset for the paper: "Attention Asymmetry in AI Layoff Discourse on X: A Computational Analysis of Capital vs Labour Amplification" Joy Bose, 2026. Code: https://gitlab.com/joyboseroy/attention-asymmetry Contents tweet_ids.csv Tweet IDs for 763 tweets collected from X (May 20-27, 2026). Columns: tweet_id, corpus_label (capital/labour), account_handle, date Per X Developer Policy, only tweet IDs are… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/ai-layoff-discourse-amplification.1 likes23 downloads4mo agoHugging Face25Kyle1668 /labeled_alignment_discourse_v1tabular1K<n<10K0 likes22 downloads10mo agoHugging Face26geodesic-research /discourse-grounded-misalignment-synthetic-scenario-datagated Synthetic Pretraining Documents for Alignment We generate synthetic pretraining documents to study how AI discourse in training data affects model alignment. Our findings suggest that natural levels of AI discourse influence model behavior; to maximally elicit this phenomenon, we upsample highly-targeted synthetic discourse covering the topics in our alignment evaluations. Overview For each question in the Articles split of our evaluation suite, we generate multiple… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-synthetic-scenario-data.text10M<n<100M2 likes20 downloads9mo agoHugging Face27Rohan1103 /pipeline-observability-discourse-forumstabular1K<n<10K0 likes19 downloads6mo agoHugging Face28mteb /hindi_discoursetext1K<n<10K0 likes18 downloads1y agoHugging Face29sileod /discourse_marker_qaDiscourse marker/connective prediction as multiple choice questions based on the Discovery datasettextquestion-answeringn<1K3 likes17 downloads4y agoHugging Face30Matthewfung /weibo_unemployment_discoursetext10K<n<100K1 likes17 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.