CoolFace
20 results

discourse

multilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes12k downloads1y agoHugging FaceCompassioninMachineLearning /caml-animal-discourse-2020-present Reddit Animal-Discourse Corpus — CLEANED (2020–present) Submissions and comments from animal-relevant subreddits, gathered via PullPush.io, covering January 2020 to the present. Built as part of research on AI-mediated value lock-in in human animal-welfare discourse. Coverage Subreddit Submissions Comments Date range (submissions) r/AnimalRights 15,719 34,686 2020-01-01 → 2025-05-19 r/AntiVegan 17,252 182,890 2020-01-01 → 2025-05-19 r/AskVegans 4… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/caml-animal-discourse-2020-present.tabulartext-classification1M<n<10M0 likes811 downloads3mo agoHugging FaceCompassioninMachineLearning /reddit-control-discourse-2016-present-pretau Reddit Control Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-control-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,456,560 of 5,930,785 records (75.1%) from 1,030,104 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-control-discourse-2016-present-pretau.tabular1M<n<10M0 likes260 downloads3mo agoHugging Facegoogle-research-datasets /coarse_discourse Dataset Card for "coarse_discourse" Dataset Summary A large corpus of discourse annotations and relations on ~10K forum threads. We collect and release a corpus of over 9,000 threads comprising over 100,000 comments manually annotated via paid crowdsourcing with discourse acts and randomly sampled from the site Reddit. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/coarse_discourse.texttext-classification100K<n<1M6 likes242 downloads3y agoHugging FaceCompassioninMachineLearning /reddit-animal-discourse-2016-present-pretau Reddit Animal Discourse 2016-present — Pre-ChatGPT Participants Subset of CompassioninMachineLearning/reddit-animal-discourse-2016-present restricted to hashed authors whose first comment in the corpus predates ChatGPT (2022-11-30). Robustness arm: isolates established human participants from the post-2022 LLM-bot / karma-farm wave. Kept 4,852,036 of 6,221,220 records (78.0%) from 343,756 pre-ChatGPT authors. author_first_seen.parquet maps every hashed author to first-seen date… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present-pretau.tabular1M<n<10M0 likes235 downloads3mo agoHugging FaceCompassioninMachineLearning /reddit-animal-discourse-2016-present Reddit Animal Discourse (2016-present) Treatment arm of the value-lock-in study, extended back to Jan 2016 to give a long pre-ChatGPT baseline for event-study leads/lags (parallel-trends test) and in-time placebo breakpoints (2017/2018/2019). 2020-present is the authoritative clean+dedup corpus; 2016-2019 is a 2,500/month-capped backfill, cleaned and deduped to the same rule. Authors salted-hashed. Pairs with the other arm for difference-in-differences / event-study analysis… See the full description on the dataset page: https://huggingface.co/datasets/CompassioninMachineLearning/reddit-animal-discourse-2016-present.tabular1M<n<10M0 likes177 downloads3mo agoHugging Face