datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.bluesky
Bluesky posts
Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders.
[!NOTE]
This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset.
Dataset Details
Size: Approximately 9 million posts
Format: JSON Lines (.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/bluesky.bluesky-posts
8 Million Bluesky Social Posts Collection
I've collected and curated 8 million public posts from Bluesky Social between November 27 - December 1, 2024, with an additional 12 million posts coming in the upcoming weeks. This growing dataset aims to provide researchers and developers with a comprehensive sample of real world social media data for analysis and experimentation. This collection represents one of the largest publicly available Bluesky datasets, offering unique insights… See the full description on the dataset page: https://huggingface.co/datasets/withalim/bluesky-posts.bluesky
Bluesky posts
Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders.
[!NOTE]
This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset.
Dataset Details
Size: Approximately 9 million posts
Format: JSON Lines (.jsonl)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky.two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/two-million-bluesky-posts.bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.bluesky-alt-text
Bluesky Alt Text: Frozen April 2026 Snapshot
Frozen historical snapshot. These files were collected in April 2026 and
will not be extended into a longitudinal series. The 279K-row corpus was
deliberately selected from accounts with high alt-text adoption, so it must
not be interpreted as a representative platform adoption estimate. The
14.5-hour Jetstream file is a short observed window, not a durable census.
Rows contain author handles and DIDs, post URIs and CIDs, post… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text.bluesky-sentiment
Bluesky Sentiment Dataset Card
Overview
Bluesky Sentiment contains posts from the agentlans/bluesky dataset, annotated for six emotions:
happiness, sadness, fear, disgust, anger, and surprise.
Annotations were generated automatically using ChatGPT, providing a nuanced, multidimensional sentiment analysis beyond simple positive/negative labels.
The dataset covers posts in multiple languages.
The few-shot config contains annotations by google/gemma-3-4b-it with 10-shot… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky-sentiment.bluesky_tpotBlueSky-ShareGPT-V0.2the bluesky sharegpt set but now filtered for only english convos via langdetect
Total conversations processed: 175163
English conversations kept: 51828
Non-English conversations removed: 123335
BlueSky-Sharegpt-V0.1https://huggingface.co/datasets/metalure/733k_bluesky_threads
Converted into sharegpt format for Finetuning. Need to filter out all non-english data. Only rows with more then 3 turns were considered to be converted.
BlueSky-Experimental-sharegptbluesky_tpot_v3TEMA-Data
TEMA-Data
Data for TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs.
Paper · Code · Models
Datasets
Component
Config
Split
Size
Temporal initialization
temporal_init
train / validation
98,401 / 500 examples
TEMA-Dialog
sft
train
40,704 dialogues / 198,195 turns
RL training
rl_schedule
train
Download
TEMA-Bench
benchmark
test
253 dialogues / 1,239 questions
Tasks
The 18-subtask guide maps the… See the full description on the dataset page: https://huggingface.co/datasets/bluesky7/TEMA-Data.bluesky_tpot_v2bluesky_engagement_posts_ktoBluesky-Sharegpt-V0.3deduped, filtered for politics (because that's shit data) and also cleared out invalid convos bringing the size down
bluesky-sentiment-sampled-classificationbluesky-pds-docs
Converted to .JSONL with this script: https://github.com/dapper-11/docs-to-json
Intended use case(s):
For training / fine-tuning and LLM
Use in an RAG system.
Downloaded from BlueSky's website
