CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes15k downloads3y agoHugging Face02hblim /top_reddit_posts_daily Top Reddit Posts Daily Dataset Summary A continuously-updated snapshot of public Reddit discourse on AI news. Each night a GitHub Actions cron job Scrapes new submissions from a configurable list of subreddits (→ data_raw/) Classifies each post with a DistilBERT sentiment model served on Replicate (→ data_scored/) Summarises daily trends for lightweight front-end consumption (→ daily_summary/) The result is an easy-to-query, time-stamped record of Reddit sentiment that… See the full description on the dataset page: https://huggingface.co/datasets/hblim/top_reddit_posts_daily.text100K<n<1M4 likes3.7k downloads11mo agoHugging Face03julien040 /hacker-news-posts Hacker News Stories Dataset This is a dataset containing approximately 4 million stories from Hacker News, exported to a Parquet file. The dataset includes the following fields: id (int64): The unique identifier of the story. title (string): The title of the story. url (string): The URL of the story. score (int64): The score of the story. time (int64): The time the story was posted, in Unix time. comments (int64): The number of comments on the story. author (string): The… See the full description on the dataset page: https://huggingface.co/datasets/julien040/hacker-news-posts.tabular100K<n<1M8 likes1.3k downloads2mo agoHugging Face04Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes879 downloads11d agoHugging Face05farida5gaber /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.tabularquestion-answering10M<n<100M0 likes425 downloads5mo agoHugging Face06Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes217 downloads2y agoHugging Face07DarjaCore /algerian-darja-forum-posts Algerian Darja Dataset A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research. 3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average Dataset at a Glance Property Value Samples 3,209,157 Total tokens 1,193,257,847 Approx. tokens 1.193B Average tokens / sample 371.83 Language Algerian Darja Format Conversational JSON Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.texttext-generation1M<n<10M2 likes202 downloads16d agoHugging Face08BEE-spoke-data /upvoteweb-posts upvoteweb: posts Posts in upvoteweb. configs [!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in. Overview of configs: default: largely unfiltered/unprocessed original data eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.imagetext-generation10M<n<100M1 likes172 downloads9mo agoHugging Face09brusic /hacker-news-who-is-hiring-posts Context This dataset contains all first-level comments to Hacker News Who Is Hiring posts from April 2011 in various formats. All data is derived from the official Firebase API and no data cleansing has occurred with the exception for removing SEEKING FREELANCER from the start of such comments.. Who wants to be hired? and Seeking Freelancer posts are included. For privacy reasons, job seeker posts will not be included. Although the data is public, do not want to create an easily… See the full description on the dataset page: https://huggingface.co/datasets/brusic/hacker-news-who-is-hiring-posts.textn<1K1 likes152 downloads7d agoHugging Face10alessiosavi /reddit-top-posts-v0tabular1M<n<10M0 likes147 downloads1y agoHugging Face11jupo-ai /filtered-danbooru-poststabular1M<n<10M0 likes121 downloads4mo agoHugging Face12dansbecker /hackernews_hiring_postsThis dataset contains postings and comments from the following recurring threads on Hacker News Ask HN: Who is hiring? Ask HN: Who wants to be hired? Freelancer? Seeking freelancer? These post types are stored in datasets called hiring, wants_to_be_hired and freelancer respectively. Each type of posting has occurred on a regular basis for several years. You can identify when each comment/listing was added through the CommentTime field. The ParentTitle also indicates the date of the parent… See the full description on the dataset page: https://huggingface.co/datasets/dansbecker/hackernews_hiring_posts.text100K<n<1M0 likes104 downloads5y agoHugging Face13YuvrajSingh9886 /reddit-posts-summarization-grpo GRPO Summarization Eval Rollouts Evaluation artifacts for all GRPO summarization checkpoints from smolcluster — a distributed GRPO training framework for Apple Silicon Mac clusters. Two base models were fine-tuned across two training strategies and six reward configurations each, then evaluated on 200 examples from the mlabonne/smoltldr test split. Judge: gpt-5-mini-2025-08-07 · Framework: DeepEval G-Eval · Rounds: 5 averaged · Metrics (each 0–1): Faithfulness · Coverage ·… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/reddit-posts-summarization-grpo.tabularsummarizationn<1K1 likes84 downloads8d agoHugging Face14KFUPM-JRCAI /arabic-generated-social-media-posts Arabic Machine-Generated Social Media Posts Dataset This dataset contains machine-generated Arabic social media posts using a text polishing approach across multiple Large Language Model (LLMs). It was created as part of the research paper: "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599). 📋 Dataset Overview The dataset addresses the need for… See the full description on the dataset page: https://huggingface.co/datasets/KFUPM-JRCAI/arabic-generated-social-media-posts.text1K<n<10K1 likes68 downloads4mo agoHugging Face15pszemraj /LocalLLaMA-posts r/LocalLLaMA posts Posts from r/LocalLLaMA pulled up through Tue Mar 3 9PM EST 2026 with arctic-shift. Now you can check if your wonderfully thought out post hasn't already been asked 30x Usage For simple semantic search, try loading it in the vectorsearch-hub-datasets space: tabulartext-generation100K<n<1M0 likes57 downloads7mo agoHugging Face16aidn /linkedin-posts-score LinkedIn Corporate Nonsense Score Dataset Ein automatisch wachsender Datensatz realer LinkedIn-Posts, bewertet nach ihrem Grad an Corporate Nonsense — gesammelt über die LinkedIn Translator App. Dataset Details Beschreibung Nutzer der App geben LinkedIn-Posts ein um sie auf ihren semantischen Kern zu reduzieren. Jeder Post wird dabei von Llama 4 Maverick automatisch anhand von 5 Metriken bewertet. Die Bewertungen und der vollständige Post-Text… See the full description on the dataset page: https://huggingface.co/datasets/aidn/linkedin-posts-score.tabulartext-classificationn<1K1 likes57 downloads2mo agoHugging Face17fdaudens /hf-blog-postsAll the Hugging Face blog posts until May 12, 2024. Includes URLs, headlines, dates, authors, and texts. textn<1K5 likes53 downloads2y agoHugging Face18fdaudens /hf-blog-posts-dpo_raw Dataset Card for hf-blog-posts-dpo_raw This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/fdaudens/hf-blog-posts-dpo_raw/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/hf-blog-posts-dpo_raw.textn<1K2 likes48 downloads2y agoHugging Face19hoaj /fb-housing-poststext1K<n<10K0 likes38 downloads2y agoHugging Face20qugemingzi /moltbook-ai-agent-posts Moltbook AI Agent Posts Dataset This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns. Dataset Statistics Total Posts: 25,445 Unique Authors: 9,955 Date Range: N/A to N/A Dataset Structure Each example contains: id: Unique post identifier title: Post title content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.tabulartext-generation10K<n<100K0 likes38 downloads8mo agoHugging Face21fernandals /UFRN-posts Dataset Card for "UFRN-posts" A base contém posts relacionados à UFRN em PT-BR. More Information needed text10K<n<100K0 likes37 downloads3y agoHugging Face22theblackcat102 /crossvalidated-posts Cross Validated / stats.stackexchange.com Dataset Summary This dataset contains all posts submitted to stats.stackexchange.com before the 30th of August 2023 formatted as Markdown text. The data is sourced from Internet Archive StackExchange Data Dump and follows the format by mikex86/stackoverflow-posts Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved due to… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/crossvalidated-posts.textquestion-answering100K<n<1M0 likes37 downloads3y agoHugging Face23roshbeed /ai-residency-hn-poststabular100K<n<1M0 likes36 downloads26d agoHugging Face24julep-ai /openai-community-posts OpenAI Community Posts This dataset is curated from the posts of the OpenAI Community Forum (https://community.openai.com). Dataset Details Dataset Description The OpenAI Community Posts dataset comprises discussions, posts, and metadata from the OpenAI Community Forum. It includes details such as discussion titles, tags, views, reply counts, post content, sentiment scores, vector embeddings for content analysis, and identifiers linking posts to… See the full description on the dataset page: https://huggingface.co/datasets/julep-ai/openai-community-posts.tabular10K<n<100K17 likes34 downloads3y agoHugging Face25theblackcat102 /datascience-stackexchange-posts Dataset Card for "datascience-stackexchange-posts" More Information needed text10K<n<100K3 likes30 downloads3y agoHugging Face26arimalabs /2.3-million-bluesky-posts2.3 Million Bluesky Posts Curated by: Gion Language(s) (NLP): Multiple (primarily English) License: Dataset usage is subject to Bluesky's Terms of Service text1M<n<10M5 likes29 downloads2y agoHugging Face27konsman /stackoverflow-posts-mobile-development-tagtabularn<1K0 likes26 downloads2y agoHugging Face28fdaudens /hf-blog-posts-splittextn<1K1 likes26 downloads2y agoHugging Face29Denn231 /only_ticker_single_posts_v_0.2tabular10K<n<100K0 likes23 downloads2y agoHugging Face30werty1248 /Korea-Related-Reddit-posts Comments: werty1248/Korea-Related-Reddit-comments This dataset may contain aggressive content. Source Subreddit comments/submissions 2005-06 to 2024-12 Korea/Korean related posts only Subset Targeted subreddits r/hanguk r/Korea* r/Korean* r/kpop* Posts Contains at least one Korean character in the selftext (body). Comments Only comments that reply to Korean posts as defined above. Statistics Total data: 3.28TB Korea/Korean-related subreddit… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Korea-Related-Reddit-posts.tabular10K<n<100K5 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.