datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paloma_subredditsreddit-subreddits25
Reddit Subreddits25 Archive
This repository contains a compressed Reddit archive organized by subreddit and record type.
It was prepared from a local bundle named reddit_torrent_download_bundle; that bundle did
not include provenance or licensing documentation, so users must independently verify that
their use complies with applicable terms, licenses, privacy requirements, and laws.
Contents
Files: 79,955 (39,974 comment shards and 39,981 submission shards)… See the full description on the dataset page: https://huggingface.co/datasets/yikeee/reddit-subreddits25.aigirlcock-subreddits-v1subreddits
Dataset Card for "subreddits"
More Information needed
cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__xlm-roberta-basecond_ft_subreddit_on_reddit__prcnt_100__test_run_False__bert-base-uncasedreddit-subreddit-nsfw-classification
Reddit Subreddit NSFW Classification
A subreddit-level NSFW / SFW / borderline classification of 67,120 subreddits,
built for responsible, ethical, and reproducible social-science research on Reddit.
This dataset is released as part of the Accelerating Social Science with Agents
and Responsible Research Using Reddit initiative. The goal of the initiative is
to leverage Reddit and other social datasets for responsible and ethical social
science, and we will release a series of… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/reddit-subreddit-nsfw-classification.cond_ft_subreddit_on_reddit__prcnt_na__test_run_True__bert-base-uncasedchange-my-view-subreddit-cleaned
Opinionated LLM
cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__roberta-basecond_ft_subreddit_on_reddit__prcnt_na__test_run_True__roberta-basecond_ft_subreddit_on_reddit__prcnt_na__test_run_Truereddit-subreddits
reddit-subreddits
One row per subreddit, derived from the January 2025 subreddit metadata files in
Binxk/pushshift-reddit.
Contents
source_type
subreddits
public
2,776,279
restricted
1,923,526
private
182,045
other
100
Total: 4,881,950 subreddits.
Null-metadata rows: about 1.5 million public rows have null subscribers,
over18, quarantine, and lang together. These are subreddits that existed at
some point but were banned, deleted, or… See the full description on the dataset page: https://huggingface.co/datasets/Binxk/reddit-subreddits.2024-election-subreddit-threads-173k
About
This dataset contains threads from 23 political subreddits from July 2024 - November 2024 (about a week after the US election).
Use this dataset as a baseline for subsets pertaining to Reddit's opinion on the 2024 election. We recommend using each thread's metadata as guidance.
E.g.,
r/politics subset
controversial comments subset
highly upvoted posts subset
leftist/liberal threads subset
etc.
Subreddits
These are the subreddits scraped. Each conversation's… See the full description on the dataset page: https://huggingface.co/datasets/BinghamtonUniversity/2024-election-subreddit-threads-173k.subreddit-postsDataset of titles of the top 1000 posts from the top 250 subreddits scraped using PRAW.
For steps to create the dataset check out the dataset script in the GitHub repo.
cond_ft_subreddit_on_reddit__prcnt_20__test_run_False__xlm-roberta-base2024-election-subreddit-threads-173kwikipedia_subredditssubreddits-v0the-antiwork-subreddit-datasetThis dataset follows the notorious subreddit /r/Antiwork, a place for many Redditors to share resources and discuss grievances with the current labour market.reddit-subreddit-scraper-sample-data
Reddit Subreddit Scraper
Scrape posts from any subreddit - title, author, score, comments, flair, text and timestamps. Run it on a schedule for social listening, brand monitoring, lead generation or market research.
What the actor scrapes
👽 Reddit Subreddit Scraper — Scrape Reddit Posts, Scores & Comments Scrape posts from any subreddit on Reddit — title, author, score, comment count, flair and full self text — and export them to JSON, CSV or Excel. This Reddit… See the full description on the dataset page: https://huggingface.co/datasets/logiover/reddit-subreddit-scraper-sample-data.StackStar_subredditsautotrain-data-pegasus-subreddit-comments-summarizer
AutoTrain Dataset for project: pegasus-subreddit-comments-summarizer
Dataset Description
This dataset has been automatically processed by AutoTrain for project pegasus-subreddit-comments-summarizer.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "I go through this every single year. We have an Ironman competition that is 2 miles… See the full description on the dataset page: https://huggingface.co/datasets/stevied67/autotrain-data-pegasus-subreddit-comments-summarizer.Confession-Subreddit-Top500This dataset was prepared by taking into account the 500 most popular posts of all time in the confession subreddit and the comments with the most votes on these posts.
subreddit_data
Dataset Card for "subreddit_data"
More Information needed
synthetic_subreddit_multiturnPortuguese-Subredditsreddit_comments_subreddit_canadareddit_subreddits_sharegptThis dataset contains 2.7M comments from various 100 subreddits.
reddit-subreddits
