CoolFace
Datasetpublic

BinghamtonUniversity/2024-election-subreddit-threads-173k

About This dataset contains threads from 23 political subreddits from July 2024 - November 2024 (about a week after the US election). Use this dataset as a baseline for subsets pertaining to Reddit's opinion on the 2024 election. We recommend using each thread's metadata as guidance. E.g., r/politics subset controversial comments subset highly upvoted posts subset leftist/liberal threads subset etc. Subreddits These are the subreddits scraped. Each… See the full description on the dataset page: https://huggingface.co/datasets/BinghamtonUniversity/2024-election-subreddit-threads-173k.

sourceHugging Faceupdated 2y agoView on Hugging Face
2likes30downloads
Dataset Card

About

This dataset contains threads from 23 political subreddits from July 2024 - November 2024 (about a week after the US election).

Use this dataset as a baseline for subsets pertaining to Reddit's opinion on the 2024 election. We recommend using each thread's metadata as guidance. E.g.,

  • —r/politics subset
  • —controversial comments subset
  • —highly upvoted posts subset
  • —leftist/liberal threads subset

etc.

Subreddits

These are the subreddits scraped. Each conversation's subreddit can be found in metadata.subreddit.name

['destiny' 'hasan_piker' 'politics' 'vaushv' 'millenials' 'news' 'worldnews' 'economics' 'socialism' 'conservative' 'libertarian' 'neoliberal' 'republican' 'democrats' 'progressive' 'daverubin' 'jordanpeterson' 'samharris' 'joerogan' 'thedavidpakmanshow' 'benshapiro' 'themajorityreport' 'seculartalk']


dataset_info: features:

  • —name: conversations list:
  • —name: content dtype: string
  • —name: role dtype: string
  • —name: metadata struct:
  • —name: controversiality dtype: int64
  • —name: normalized_controversiality dtype: float64
  • —name: post struct:
  • —name: author dtype: string
  • —name: downvotes dtype: int64
  • —name: flair dtype: string
  • —name: score dtype: int64
  • —name: suggested_sort dtype: string
  • —name: upvote_ratio dtype: float64
  • —name: upvotes dtype: int64
  • —name: subreddit struct:
  • —name: name dtype: string
  • —name: subscribers dtype: int64 splits:
  • —name: train numbytes: 201931627 numexamples: 173583 downloadsize: 117663856 datasetsize: 201931627 configs:
  • —configname: default datafiles:
  • —split: train path: data/train-* license: mit language:
  • —en tags:
  • —not-for-all-audiences
  • —reddit prettyname: r sizecategories:
  • —100K<n<1M ---