datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.hacker-news-posts
Hacker News Stories Dataset
This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields:
id (int64): The unique identifier of the story.
title (string): The title of the story.
url (string): The URL of the story.
score (int64): The score of the story.
time (int64): The time the story was posted, in Unix time.
comments (int64): The number of comments on the story.
author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.donald-trump-truth-social-posts
Donald Trump Truth Social Posts Archive
Archive overview
36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables.
It also includes streamable image media plus video metadata and transcripts where the source provides them.
The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
stackoverflow-posts
StackOverflow Posts Markdown
Dataset Summary
This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text.
The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text.
The data is sourced from Internet Archive StackExchange Data Dump.
Dataset Structure
Each record corresponds to one post of a particular type.
Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.CSSR-S_labelled_suicidewatch_posts_reddit
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
Full code and supplementary materials are available at https://github.com/av9ash/llm_cssrs_code.
License and Citation
This project is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.Any use or reuse of this work please cite the following:
@article{patil2025evaluating,
title={Evaluating Reasoning LLMs for Suicide Screening with the… See the full description on the dataset page: https://huggingface.co/datasets/av9ash/CSSR-S_labelled_suicidewatch_posts_reddit.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.upvoteweb-posts
upvoteweb: posts
Posts in upvoteweb.
configs
[!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in.
Overview of configs:
default: largely unfiltered/unprocessed original data
eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier
en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.reddit-top-posts-v0moltbook_postsMoltbook Posts Dataset
~Posts scraped from Moltbook's public API (no auth required). AI agent social network discussions.
Features:
title (str)
content (str)
created_at (datetime)
author_name (str)
submolt_display_name (str)
upvotes (int)
comment_count (int)
License: MIT – public data.
Citation: Moltbook public API – https://www.moltbook.com
filtered-danbooru-postsreddit_finance_posts_sp500
Reddit Finance Posts Dataset SP500
This dataset contains 431,923 Reddit posts collected from 20 finance-related subreddits via the Reddit API.
As keywords, all S&P 500 companies were used. The included subreddits are:
stocks, wallstreetbets, investing, StockMarket, options, RobinHood,
pennystocks, SecurityAnalysis, personalfinance, Dividends, CryptoCurrency,
CryptoMarkets, ETFs, FinancialIndependence, ValueInvesting, quant,
algotrading, forex, economy, Superstonk, spacs… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_finance_posts_sp500.reddit-posts-summarization-grpo
GRPO Summarization Eval Rollouts
Evaluation artifacts for all GRPO summarization checkpoints from smolcluster — a distributed GRPO training framework for Apple Silicon Mac clusters.
Two base models were fine-tuned across two training strategies and six reward configurations each, then evaluated on 200 examples from the mlabonne/smoltldr test split.
Judge: gpt-5-mini-2025-08-07 · Framework: DeepEval G-Eval · Rounds: 5 averaged · Metrics (each 0–1): Faithfulness · Coverage ·… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/reddit-posts-summarization-grpo.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
LocalLLaMA-posts
r/LocalLLaMA posts
Posts from r/LocalLLaMA pulled up through Tue Mar 3 9PM EST 2026 with arctic-shift. Now you can check if your wonderfully thought out post hasn't already been asked 30x
Usage
For simple semantic search, try loading it in the vectorsearch-hub-datasets space:
linkedin-posts-score
LinkedIn Corporate Nonsense Score Dataset
Ein automatisch wachsender Datensatz realer LinkedIn-Posts, bewertet nach ihrem Grad an Corporate Nonsense — gesammelt über die LinkedIn Translator App.
Dataset Details
Beschreibung
Nutzer der App geben LinkedIn-Posts ein um sie auf ihren semantischen Kern zu reduzieren. Jeder Post wird dabei von Llama 4 Maverick automatisch anhand von 5 Metriken bewertet. Die Bewertungen und der vollständige Post-Text… See the full description on the dataset page: https://huggingface.co/datasets/aidn/linkedin-posts-score.linkedin_postsreddit_nosleep_postsreddit_finance_posts_apple-tesla-microsoft
Reddit Finance Posts Dataset (Apple, Tesla, Microsoft)
This dataset contains 12046 Reddit Posts collected from 20 finance-related subreddits via the Reddit API using the keywords Apple, Tesla, and Microsoft.
The included subreddits are:
stocks, wallstreetbets, investing, StockMarket, options, RobinHood,
pennystocks, SecurityAnalysis, personalfinance, Dividends, CryptoCurrency,
CryptoMarkets, ETFs, FinancialIndependence, ValueInvesting, quant,
algotrading, forex, economy, Superstonk… See the full description on the dataset page: https://huggingface.co/datasets/emilpartow/reddit_finance_posts_apple-tesla-microsoft.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
moltbook-ai-agent-posts
Moltbook AI Agent Posts Dataset
This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns.
Dataset Statistics
Total Posts: 25,445
Unique Authors: 9,955
Date Range: N/A to N/A
Dataset Structure
Each example contains:
id: Unique post identifier
title: Post title
content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.ai-residency-hn-postsopenai-community-posts
OpenAI Community Posts
This dataset is curated from the posts of the OpenAI Community Forum (https://community.openai.com).
Dataset Details
Dataset Description
The OpenAI Community Posts dataset comprises discussions, posts, and metadata from the OpenAI Community Forum.
It includes details such as discussion titles, tags, views, reply counts, post content, sentiment scores, vector embeddings for content analysis, and identifiers linking posts to… See the full description on the dataset page: https://huggingface.co/datasets/julep-ai/openai-community-posts.hf-posts
Hugging Face Posts
This dataset contains posts scraped from https://huggingface.co/posts.
It includes all posts published from the launch date on December 23, 2023, up to November 24, 2024, at 15:40.
sentiment_postsstackoverflow-posts-mobile-development-tagonly_ticker_single_posts_v_0.2Korea-Related-Reddit-posts
Comments: werty1248/Korea-Related-Reddit-comments
This dataset may contain aggressive content.
Source
Subreddit comments/submissions 2005-06 to 2024-12
Korea/Korean related posts only
Subset
Targeted subreddits
r/hanguk
r/Korea*
r/Korean*
r/kpop*
Posts
Contains at least one Korean character in the selftext (body).
Comments
Only comments that reply to Korean posts as defined above.
Statistics
Total data: 3.28TB
Korea/Korean-related subreddit… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Korea-Related-Reddit-posts.scraped-forum-postsreddit_posts
Dataset Card for "reddit_posts"
More Information needed
