datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.eevblog-posts
🛠️ EEVblog Forum Dataset: The Electronics Mentor
Stop training on synthetic data. Train on real engineering wisdom. 200K+ authentic technical conversations where beginners learn from seasoned engineers, troubleshooting experts guide newcomers, and practical wisdom gets passed down through generations of makers.
🚀 What Makes This Special?
This isn't just another Q&A dataset. This is 200,756 posts of authentic mentor-apprentice dialogue where beginners learn from seasoned… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/eevblog-posts.hackaday-posts
🚀 Hackaday Universe: 50K+ Tech Articles & Vibrant Maker Conversations
Dive into the ultimate collection of Hackaday's tech universe! This isn't just another dataset—it's a living archive of maker culture, featuring 54,599+ articles with complete comment threads where brilliant minds collide, debate, and innovate together.
🔥 Why This Dataset Rocks
🤖 Perfect for AI Training
Train models on authentic technical writing and community interactions
Learn from real… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/hackaday-posts.russian-smm-posts
zxc0zxc0zxc/russian-smm-posts
A small Russian-language dataset of social media writing examples based on posts from major Russian media Telegram
channels.
The dataset contains up to 2,000 examples collected via Telegram Channel Export and then automatically reformatted,
structured, and annotated with ChatGPT 5.3. In practice, this dataset was created through a distillation-style pipeline,
where source posts were converted into chat-style supervised fine-tuning examples with system… See the full description on the dataset page: https://huggingface.co/datasets/zxc0zxc0zxc/russian-smm-posts.vericava-posts-2026
vericava-posts-2026
Cleaned data of posts from @vericava.
bordaru-posts
Dataset Card for Borda.ru Posts
Dataset Summary
This dataset contains posts scraped from Borda.ru, a Russian platform for hosting various discussion forums on a wide range of topics. Each entry in the dataset represents a post from the website, including its content, author, URL, and other relevant information. The dataset contains 5,251,346 unique messages. The dataset was deduplicated based on the "content" value, which removed spam and other low-quality data, keeping… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/bordaru-posts.
