CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WillHeld /paloma_subredditstext10K<n<100K0 likes213 downloads1y agoHugging Face02yikeee /reddit-subreddits25 Reddit Subreddits25 Archive This repository contains a compressed Reddit archive organized by subreddit and record type. It was prepared from a local bundle named reddit_torrent_download_bundle; that bundle did not include provenance or licensing documentation, so users must independently verify that their use complies with applicable terms, licenses, privacy requirements, and laws. Contents Files: 79,955 (39,974 comment shards and 39,981 submission shards)… See the full description on the dataset page: https://huggingface.co/datasets/yikeee/reddit-subreddits25.0 likes213 downloads14d agoHugging Face03Nicfingshelby /aigirlcock-subreddits-v1imagen<1K2 likes116 downloads3mo agoHugging Face04alhosseini /subreddits Dataset Card for "subreddits" More Information needed text1M<n<10M1 likes98 downloads3y agoHugging Face05emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__xlm-roberta-base1M<n<10M0 likes84 downloads4y agoHugging Face06emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__bert-base-uncased1M<n<10M0 likes52 downloads4y agoHugging Face07ssingh22 /reddit-subreddit-nsfw-classification Reddit Subreddit NSFW Classification A subreddit-level NSFW / SFW / borderline classification of 67,120 subreddits, built for responsible, ethical, and reproducible social-science research on Reddit. This dataset is released as part of the Accelerating Social Science with Agents and Responsible Research Using Reddit initiative. The goal of the initiative is to leverage Reddit and other social datasets for responsible and ethical social science, and we will release a series of… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/reddit-subreddit-nsfw-classification.text-classification10K<n<100K0 likes49 downloads1mo agoHugging Face08emilylearning /cond_ft_subreddit_on_reddit__prcnt_na__test_run_True__bert-base-uncased0 likes46 downloads4y agoHugging Face09Siddish /change-my-view-subreddit-cleaned Opinionated LLM texttext-generation1K<n<10K1 likes43 downloads3y agoHugging Face10emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__roberta-base1M<n<10M0 likes42 downloads4y agoHugging Face11emilylearning /cond_ft_subreddit_on_reddit__prcnt_na__test_run_True__roberta-base0 likes39 downloads4y agoHugging Face12emilylearning /cond_ft_subreddit_on_reddit__prcnt_na__test_run_True0 likes35 downloads4y agoHugging Face13Binxk /reddit-subredditsgated reddit-subreddits One row per subreddit, derived from the January 2025 subreddit metadata files in Binxk/pushshift-reddit. Contents source_type subreddits public 2,776,279 restricted 1,923,526 private 182,045 other 100 Total: 4,881,950 subreddits. Null-metadata rows: about 1.5 million public rows have null subscribers, over18, quarantine, and lang together. These are subreddits that existed at some point but were banned, deleted, or… See the full description on the dataset page: https://huggingface.co/datasets/Binxk/reddit-subreddits.tabular1M<n<10M1 likes32 downloads1mo agoHugging Face14BinghamtonUniversity /2024-election-subreddit-threads-173k About This dataset contains threads from 23 political subreddits from July 2024 - November 2024 (about a week after the US election). Use this dataset as a baseline for subsets pertaining to Reddit's opinion on the 2024 election. We recommend using each thread's metadata as guidance. E.g., r/politics subset controversial comments subset highly upvoted posts subset leftist/liberal threads subset etc. Subreddits These are the subreddits scraped. Each conversation's… See the full description on the dataset page: https://huggingface.co/datasets/BinghamtonUniversity/2024-election-subreddit-threads-173k.text100K<n<1M2 likes29 downloads2y agoHugging Face15daspartho /subreddit-postsDataset of titles of the top 1000 posts from the top 250 subreddits scraped using PRAW. For steps to create the dataset check out the dataset script in the GitHub repo. text100K<n<1M2 likes26 downloads4y agoHugging Face16emilylearning /cond_ft_subreddit_on_reddit__prcnt_20__test_run_False__xlm-roberta-base0 likes24 downloads4y agoHugging Face17brianmatzelle /2024-election-subreddit-threads-173ktext100K<n<1M0 likes21 downloads2y agoHugging Face18Organika /wikipedia_subredditstext1K<n<10K0 likes17 downloads3y agoHugging Face19alessiosavi /subreddits-v0tabular100K<n<1M0 likes15 downloads1y agoHugging Face20SocialGrep /the-antiwork-subreddit-datasetThis dataset follows the notorious subreddit /r/Antiwork, a place for many Redditors to share resources and discuss grievances with the current labour market.text100K<n<1M1 likes14 downloads4y agoHugging Face21logiover /reddit-subreddit-scraper-sample-data Reddit Subreddit Scraper Scrape posts from any subreddit - title, author, score, comments, flair, text and timestamps. Run it on a schedule for social listening, brand monitoring, lead generation or market research. What the actor scrapes 👽 Reddit Subreddit Scraper — Scrape Reddit Posts, Scores & Comments Scrape posts from any subreddit on Reddit — title, author, score, comment count, flair and full self text — and export them to JSON, CSV or Excel. This Reddit… See the full description on the dataset page: https://huggingface.co/datasets/logiover/reddit-subreddit-scraper-sample-data.tabularn<1K0 likes14 downloads4mo agoHugging Face22Organika /StackStar_subredditstextn<1K0 likes13 downloads2y agoHugging Face23stevied67 /autotrain-data-pegasus-subreddit-comments-summarizer AutoTrain Dataset for project: pegasus-subreddit-comments-summarizer Dataset Description This dataset has been automatically processed by AutoTrain for project pegasus-subreddit-comments-summarizer. Languages The BCP-47 code for the dataset's language is en. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "text": "I go through this every single year. We have an Ironman competition that is 2 miles… See the full description on the dataset page: https://huggingface.co/datasets/stevied67/autotrain-data-pegasus-subreddit-comments-summarizer.summarization0 likes12 downloads3y agoHugging Face24Oguzz07 /Confession-Subreddit-Top500This dataset was prepared by taking into account the 500 most popular posts of all time in the confession subreddit and the comments with the most votes on these posts. text1K<n<10K0 likes12 downloads2y agoHugging Face25Asap7772 /subreddit_data Dataset Card for "subreddit_data" More Information needed tabular1K<n<10K0 likes12 downloads3y agoHugging Face26snap-stanford /synthetic_subreddit_multiturntextn<1K0 likes9 downloads1y agoHugging Face27RexTRO111 /Portuguese-Subredditstextn<1K0 likes9 downloads2mo agoHugging Face28beenakurian /reddit_comments_subreddit_canadatexttext-classification1K<n<10K0 likes8 downloads3y agoHugging Face29adamo1139 /reddit_subreddits_sharegptThis dataset contains 2.7M comments from various 100 subreddits. text1M<n<10M2 likes8 downloads2y agoHugging Face30cochiseruhulessin /reddit-subredditstabular1K<n<10K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.