CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikex86 /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/mikex86/stackoverflow-posts.tabularquestion-answering10M<n<100M63 likes15k downloads3y agoHugging Face02Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes879 downloads11d agoHugging Face03farida5gaber /stackoverflow-posts StackOverflow Posts Markdown Dataset Summary This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved… See the full description on the dataset page: https://huggingface.co/datasets/farida5gaber/stackoverflow-posts.tabularquestion-answering10M<n<100M0 likes425 downloads5mo agoHugging Face04Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes217 downloads2y agoHugging Face05DarjaCore /algerian-darja-forum-posts Algerian Darja Dataset A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research. 3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average Dataset at a Glance Property Value Samples 3,209,157 Total tokens 1,193,257,847 Approx. tokens 1.193B Average tokens / sample 371.83 Language Algerian Darja Format Conversational JSON Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.texttext-generation1M<n<10M2 likes202 downloads16d agoHugging Face06BEE-spoke-data /upvoteweb-posts upvoteweb: posts Posts in upvoteweb. configs [!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in. Overview of configs: default: largely unfiltered/unprocessed original data eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.imagetext-generation10M<n<100M1 likes172 downloads9mo agoHugging Face07DSULT-Core /bluesky-298-million-Posts So far... 1 Million (Daniel)2 Million (Alpindale)20 Million (informatiker) Yall are weak. How about... 298 Million posts? License GAYSEX-Dont Be A Prick License What happened? Change of hearts. I've relaxed the restrictions. Just read the license instead. (It's quite hands off as long as you don't want to stir drama) text-generation50 likes118 downloads2y agoHugging Face08YuvrajSingh9886 /reddit-posts-summarization-grpo GRPO Summarization Eval Rollouts Evaluation artifacts for all GRPO summarization checkpoints from smolcluster — a distributed GRPO training framework for Apple Silicon Mac clusters. Two base models were fine-tuned across two training strategies and six reward configurations each, then evaluated on 200 examples from the mlabonne/smoltldr test split. Judge: gpt-5-mini-2025-08-07 · Framework: DeepEval G-Eval · Rounds: 5 averaged · Metrics (each 0–1): Faithfulness · Coverage ·… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/reddit-posts-summarization-grpo.tabularsummarizationn<1K1 likes84 downloads8d agoHugging Face09itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes64 downloads1mo agoHugging Face10nyuuzyou /womanru-posts Dataset Card for Woman.ru Forum Posts Dataset Summary This dataset contains 1,308,238 forum posts from Woman.ru, a popular Russian-language information and entertainment portal. Woman.ru is one of the most visited women's sites in Runet (Russian Internet). The dataset covers posts from around 2005 to 2024, providing a comprehensive view of discussions on the platform over nearly two decades. The content includes original posts and replies on various topics, offering… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/womanru-posts.texttext-generation1M<n<10M2 likes59 downloads2y agoHugging Face11pszemraj /LocalLLaMA-posts r/LocalLLaMA posts Posts from r/LocalLLaMA pulled up through Tue Mar 3 9PM EST 2026 with arctic-shift. Now you can check if your wonderfully thought out post hasn't already been asked 30x Usage For simple semantic search, try loading it in the vectorsearch-hub-datasets space: tabulartext-generation100K<n<1M0 likes57 downloads7mo agoHugging Face12qugemingzi /moltbook-ai-agent-posts Moltbook AI Agent Posts Dataset This dataset contains posts and conversations from Moltbook.com, a platform for AI character roleplay and interaction. It was collected as part of a research project comparing synthetic (AI-generated) and organic (human-generated) discourse patterns. Dataset Statistics Total Posts: 25,445 Unique Authors: 9,955 Date Range: N/A to N/A Dataset Structure Each example contains: id: Unique post identifier title: Post title content:… See the full description on the dataset page: https://huggingface.co/datasets/qugemingzi/moltbook-ai-agent-posts.tabulartext-generation10K<n<100K0 likes38 downloads8mo agoHugging Face13theblackcat102 /crossvalidated-posts Cross Validated / stats.stackexchange.com Dataset Summary This dataset contains all posts submitted to stats.stackexchange.com before the 30th of August 2023 formatted as Markdown text. The data is sourced from Internet Archive StackExchange Data Dump and follows the format by mikex86/stackoverflow-posts Dataset Structure Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved due to… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/crossvalidated-posts.textquestion-answering100K<n<1M0 likes37 downloads3y agoHugging Face14fast-flash /fast-flash-hackernews-posts Fast Flash | HackerNews Posts Dataset Exploratory Analysis Take a look at some fascinating findings from this dataset on our website. Dataset Summary We release dataset of all HackerNews posts. The dataset includes 35,316,999 posts and was collected in March 2023. You can also find a dataset of all users right here. Dataset Structure The post objects in this dataset are structured according to HackerNews' API specification. About the Author… See the full description on the dataset page: https://huggingface.co/datasets/fast-flash/fast-flash-hackernews-posts.texttext-classification10M<n<100M3 likes18 downloads3y agoHugging Face15hybridfree /eevblog-posts 🛠️ EEVblog Forum Dataset: The Electronics Mentor Stop training on synthetic data. Train on real engineering wisdom. 200K+ authentic technical conversations where beginners learn from seasoned engineers, troubleshooting experts guide newcomers, and practical wisdom gets passed down through generations of makers. 🚀 What Makes This Special? This isn't just another Q&A dataset. This is 200,756 posts of authentic mentor-apprentice dialogue where beginners learn from seasoned… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/eevblog-posts.textquestion-answering100K<n<1M0 likes18 downloads9mo agoHugging Face16danielrosehill /Blog-Poststext-generation1K<n<10K0 likes16 downloads1y agoHugging Face17hybridfree /hackaday-posts 🚀 Hackaday Universe: 50K+ Tech Articles & Vibrant Maker Conversations Dive into the ultimate collection of Hackaday's tech universe! This isn't just another dataset—it's a living archive of maker culture, featuring 54,599+ articles with complete comment threads where brilliant minds collide, debate, and innovate together. 🔥 Why This Dataset Rocks 🤖 Perfect for AI Training Train models on authentic technical writing and community interactions Learn from real… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/hackaday-posts.imagetext-generation10K<n<100K0 likes16 downloads9mo agoHugging Face18nyuuzyou /smartlab-posts Dataset Card for Smart-lab.ru Posts Dataset Summary This dataset contains posts scraped from Smart-lab.ru, a Russian platform for discussing up-to-date stock exchange information, market news, investment ideas, and trading methods. Each entry in the dataset represents a post from the website, including its title, content, author, and a unique identifier. Languages The dataset is primarily in Russian, though some posts may contain content in other languages.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/smartlab-posts.texttext-generation100K<n<1M0 likes12 downloads2y agoHugging Face19Spierocho /channel_poststexttext-generation10K<n<100K0 likes11 downloads2y agoHugging Face20zxc0zxc0zxc /russian-smm-posts zxc0zxc0zxc/russian-smm-posts A small Russian-language dataset of social media writing examples based on posts from major Russian media Telegram channels. The dataset contains up to 2,000 examples collected via Telegram Channel Export and then automatically reformatted, structured, and annotated with ChatGPT 5.3. In practice, this dataset was created through a distillation-style pipeline, where source posts were converted into chat-style supervised fine-tuning examples with system… See the full description on the dataset page: https://huggingface.co/datasets/zxc0zxc0zxc/russian-smm-posts.texttext-generation1K<n<10K0 likes11 downloads7mo agoHugging Face21menheraorg /vericava-posts-2026 vericava-posts-2026 Cleaned data of posts from @vericava. texttext-generation10K<n<100K0 likes10 downloads3mo agoHugging Face22nyuuzyou /bordaru-posts Dataset Card for Borda.ru Posts Dataset Summary This dataset contains posts scraped from Borda.ru, a Russian platform for hosting various discussion forums on a wide range of topics. Each entry in the dataset represents a post from the website, including its content, author, URL, and other relevant information. The dataset contains 5,251,346 unique messages. The dataset was deduplicated based on the "content" value, which removed spam and other low-quality data, keeping… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/bordaru-posts.texttext-generation1M<n<10M1 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.