datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
two-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.bluesky
Bluesky posts
Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders.
[!NOTE]
This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset.
Dataset Details
Size: Approximately 9 million posts
Format: JSON Lines (.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/bluesky.AI-DeepResearch-BenchReportbluesky-posts
8 Million Bluesky Social Posts Collection
I've collected and curated 8 million public posts from Bluesky Social between November 27 - December 1, 2024, with an additional 12 million posts coming in the upcoming weeks. This growing dataset aims to provide researchers and developers with a comprehensive sample of real world social media data for analysis and experimentation. This collection represents one of the largest publicly available Bluesky datasets, offering unique insights… See the full description on the dataset page: https://huggingface.co/datasets/withalim/bluesky-posts.bluesky-alt-text-observatory
Bluesky Accessibility Observatory
This is a focused longitudinal observation of declared image descriptions in
public Bluesky post commits. It begins with archive-format v2 and does not
include the biased April 2026 snapshot corpus.
daily_metrics and daily_language_metrics are aggregate observations at post
creation time. description_sample is a deterministic, uniform bottom-k
sample of non-empty descriptions after a 48-hour correction window. It uses
keyed pseudonyms, not… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text-observatory.voxceleb2-mp4-binarybluesky-embeddings-daily
🛰️ Bluesky AI Analysis: Public Post Embeddings
This dataset contains vector embeddings of public posts from the Bluesky Social network, generated for the purpose of semantic search, discovery, and language model experimentation.
📦 Contents
Each row in the dataset includes:
uri: The AT URI of the post.
created_at: The full timestamp when the post was created.
created_date: The UTC calendar date (YYYY-MM-DD).
created_hour: The UTC hour of day (0–23).
text: The post's… See the full description on the dataset page: https://huggingface.co/datasets/wildwood77/bluesky-embeddings-daily.bluesky
Bluesky posts
Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders.
[!NOTE]
This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset.
Dataset Details
Size: Approximately 9 million posts
Format: JSON Lines (.jsonl)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky.MedExQA 🧑⚕️ MedExQA
Medical Question Answering Benchmark with Multiple Explanations
📄 Paper • ⭐ Code • ⏬ Dataset • ⚕️ MedPhi2
🆕 News
[July 2024] Our work is accepted at the ACL2024 BioNLP workshop.
[July 2024] We release MedExQA dataset.
[June 2024] We release MedPhi2 and Paper.
Benchmark Summary
MedExQA is a novel benchmark in medical question-answering, to evaluate large language models’ (LLMs) understanding of medical knowledge through explanations.
The… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/MedExQA.libriheavy-cosyvoicetwo-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/two-million-bluesky-posts.bluesky-datatset@inproceedings{10.1145/3646547.3688407,
author = {Balduf, Leonhard and Sokoto, Saidu and Ascigil, Onur and Tyson, Gareth and Scheuermann, Bj\"{o}rn and Korczy\'{n}ski, Maciej and Castro, Ignacio and Kr\'{o}l, Michaundefined},
title = {Looking AT the Blue Skies of Bluesky},
year = {2024},
isbn = {9798400705922},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3646547.3688407},
doi = {10.1145/3646547.3688407},
abstract = {The… See the full description on the dataset page: https://huggingface.co/datasets/labofsahil/bluesky-datatset.bluesky-298-million-Posts
So far...
1 Million (Daniel)2 Million (Alpindale)20 Million (informatiker)
Yall are weak. How about... 298 Million posts?
License
GAYSEX-Dont Be A Prick License
What happened?
Change of hearts. I've relaxed the restrictions. Just read the license instead. (It's quite hands off as long as you don't want to stir drama)
bluesky
Bluesky User Events Stream
This repository contains user events on the Bluesky social network from its inception in Nov. 2022 to July 2023. The events captured include likes, follows, blocks, and other user interactions on the platform. More recent events will be added over time.
The data is stored in JSONL format:
{
"createdAt": "2024-07-21T15:30:00Z",
"$type": "app.bsky.feed.like",
"did": "did:plc:123456",
"uri": "at://did:plc:123456/app.bsky.feed.like/78910",
... //… See the full description on the dataset page: https://huggingface.co/datasets/hallofstairs/bluesky.bluesky
Five Million bluesky posts
This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data.
Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky.bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.bluesky-ai-discourse-corpus
Bluesky AI-Discourse Corpus
21M+ posts retrieved via AI-related keyword search using Bluesky's
public search API (app.bsky.feed.searchPosts).
Built for AI-perception/sentiment research: what people say about AI
tools, companies, models, art, coding, safety, and each other — including
the pro-AI/anti-AI contrast communities.
What's inside
21 Million posts (deduplicated by post URI)
Collected via 141 taxonomy queries: 66 category chunks across 13
topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.chemical_language_understanding_benchmark🧪🔋 Chemical Language Understanding Benchmark 🛢️🧴
Benchmark Summary
Chemistry Language Understanding Benchmark is published in ACL2023 industry track to facilitate NLP research in chemical industry ACL2023 Industry Track.
From our understanding, it is one of the first benchmark datasets with tasks for both patent and literature articles provided by the industrial organization.
All the datasets are annotated by professional chemists.
Languages
The language of this benchmark is English.… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/chemical_language_understanding_benchmark.three-million-bluesky
3 million bluesky posts
raw.zip is the ~5m posts that i initially pulled that were full of duplicates, where i removed the duplicates and compiled those into 2.6 million posts in final_posts.jsonl
the rest of the .jsonl files in the data folder, and not inside the raw.zip are unchecked and may have duplicates. feel free to write your own script to remove them, but this should contain around 3 million unique posts.
AI-Essaysynthetic_discharge_summ
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is a subset of the dataset for Asclepius model (arxiv).
The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5.
We filtered the summarization task for discharge notes. The dataset contains 13,584 notes.
Supported Tasks
This dataset covers below summarization task
Languages
English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.bluesky-alt-text
Bluesky Alt Text: Frozen April 2026 Snapshot
Frozen historical snapshot. These files were collected in April 2026 and
will not be extended into a longitudinal series. The 279K-row corpus was
deliberately selected from accounts with high alt-text adoption, so it must
not be interpreted as a representative platform adoption estimate. The
14.5-hour Jetstream file is a short observed window, not a durable census.
Rows contain author handles and DIDs, post URIs and CIDs, post… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text.2.3-million-bluesky-posts2.3 Million Bluesky Posts
Curated by: Gion
Language(s) (NLP): Multiple (primarily English)
License: Dataset usage is subject to Bluesky's Terms of Service
bluesky-ten-million
Ten Million bluesky posts
This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data.
Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky-ten-million.bluesky-sentiment
Bluesky Sentiment Dataset Card
Overview
Bluesky Sentiment contains posts from the agentlans/bluesky dataset, annotated for six emotions:
happiness, sadness, fear, disgust, anger, and surprise.
Annotations were generated automatically using ChatGPT, providing a nuanced, multidimensional sentiment analysis beyond simple positive/negative labels.
The dataset covers posts in multiple languages.
The few-shot config contains annotations by google/gemma-3-4b-it with 10-shot… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky-sentiment.bluesky_tpotbluesky_profiles
Bluesky Network (Profiles and Follows)
This is a scraped mirror of the Bluesky (https://bsky.app/) social graph. It includes profile information (did, handle, display name, indexed at, follows count, followers count, posts count, and descriptions). The follow graph is (did, did) relationships, with created at timestamp. There is also a calculated PageRank of the follows graph.
Notes:
Consult the Bluesky / AT Proto API docs for explainations for fields.
Scraping prioritizes larger… See the full description on the dataset page: https://huggingface.co/datasets/andrewconner/bluesky_profiles.gender-bluesky-classification-v3This is a dataset that contains 1 million text entries with their corresponding gender labels for gender classification.
this is the training set's gender distribution:
this is the test set's gender distribution:
This dataset was obtained by scraping 1 million messages from the Bluesky Firehose.
here is how the program i made works.
it would check each message in the firehose for the posters did and go check the poster's profile bio for pronouns such as "she/her" or "he/him".
If pronouns… See the full description on the dataset page: https://huggingface.co/datasets/breadlicker45/gender-bluesky-classification-v3.Bluesky1bluesky-five-million
Five Million bluesky posts
This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data.
Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky-five-million.
