bluesky
Datasets
All datasets matching “bluesky”two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.bluesky
Bluesky posts
Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders.
[!NOTE]
This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset.
Dataset Details
Size: Approximately 9 million posts
Format: JSON Lines (.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/bluesky.AI-DeepResearch-BenchReportbluesky-posts
8 Million Bluesky Social Posts Collection
I've collected and curated 8 million public posts from Bluesky Social between November 27 - December 1, 2024, with an additional 12 million posts coming in the upcoming weeks. This growing dataset aims to provide researchers and developers with a comprehensive sample of real world social media data for analysis and experimentation. This collection represents one of the largest publicly available Bluesky datasets, offering unique insights… See the full description on the dataset page: https://huggingface.co/datasets/withalim/bluesky-posts.bluesky-alt-text-observatory
Bluesky Accessibility Observatory
This is a focused longitudinal observation of declared image descriptions in
public Bluesky post commits. It begins with archive-format v2 and does not
include the biased April 2026 snapshot corpus.
daily_metrics and daily_language_metrics are aggregate observations at post
creation time. description_sample is a deterministic, uniform bottom-k
sample of non-empty descriptions after a 48-hour correction window. It uses
keyed pseudonyms, not… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text-observatory.voxceleb2-mp4-binary
