CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alpindale /two-million-bluesky-posts 2 Million Bluesky Posts This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data. The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model. Dataset Details Dataset Description This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.text1M<n<10M206 likes1.4k downloads2y agoHugging Face02bobHe2099 /bluesky Bluesky posts Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders. [!NOTE] This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset. Dataset Details Size: Approximately 9 million posts Format: JSON Lines (.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/bluesky.text10M<n<100M0 likes395 downloads2mo agoHugging Face03BlueSkyXN /AI-DeepResearch-BenchReport3 likes380 downloads6mo agoHugging Face04withalim /bluesky-posts 8 Million Bluesky Social Posts Collection I've collected and curated 8 million public posts from Bluesky Social between November 27 - December 1, 2024, with an additional 12 million posts coming in the upcoming weeks. This growing dataset aims to provide researchers and developers with a comprehensive sample of real world social media data for analysis and experimentation. This collection represents one of the largest publicly available Bluesky datasets, offering unique insights… See the full description on the dataset page: https://huggingface.co/datasets/withalim/bluesky-posts.text1M<n<10M5 likes351 downloads2y agoHugging Face05lukeslp /bluesky-alt-text-observatory Bluesky Accessibility Observatory This is a focused longitudinal observation of declared image descriptions in public Bluesky post commits. It begins with archive-format v2 and does not include the biased April 2026 snapshot corpus. daily_metrics and daily_language_metrics are aggregate observations at post creation time. description_sample is a deterministic, uniform bottom-k sample of non-empty descriptions after a 48-hour correction window. It uses keyed pseudonyms, not… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text-observatory.tabular10K<n<100K0 likes329 downloads8d agoHugging Face06blueskyheaven /voxceleb2-mp4-binarytext1M<n<10M0 likes306 downloads1y agoHugging Face07wildwood77 /bluesky-embeddings-daily 🛰️ Bluesky AI Analysis: Public Post Embeddings This dataset contains vector embeddings of public posts from the Bluesky Social network, generated for the purpose of semantic search, discovery, and language model experimentation. 📦 Contents Each row in the dataset includes: uri: The AT URI of the post. created_at: The full timestamp when the post was created. created_date: The UTC calendar date (YYYY-MM-DD). created_hour: The UTC hour of day (0–23). text: The post's… See the full description on the dataset page: https://huggingface.co/datasets/wildwood77/bluesky-embeddings-daily.textfeature-extraction10M<n<100M1 likes290 downloads1y agoHugging Face08agentlans /bluesky Bluesky posts Approximately 9 million public Bluesky posts, processed and cleaned for machine learning research and experimentation. The dataset has been normalized and filtered to remove duplicates, with sensitive information replaced by placeholders. [!NOTE] This dataset isn't directly from Bluesky itself. It's a processed version of the Roronotalt/bluesky-ten-million dataset. Dataset Details Size: Approximately 9 million posts Format: JSON Lines (.jsonl) Source:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky.text10M<n<100M2 likes276 downloads2y agoHugging Face09bluesky333 /MedExQA 🧑‍⚕️ MedExQA Medical Question Answering Benchmark with Multiple Explanations 📄 Paper • ⭐ Code • ⏬ Dataset • ⚕️ MedPhi2 🆕 News [July 2024] Our work is accepted at the ACL2024 BioNLP workshop. [July 2024] We release MedExQA dataset. [June 2024] We release MedPhi2 and Paper. Benchmark Summary MedExQA is a novel benchmark in medical question-answering, to evaluate large language models’ (LLMs) understanding of medical knowledge through explanations. The… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/MedExQA.multiple-choicen<1K7 likes257 downloads2y agoHugging Face10blueskyheaven /libriheavy-cosyvoice0 likes191 downloads9mo agoHugging Face11bobHe2099 /two-million-bluesky-posts 2 Million Bluesky Posts This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data. The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model. Dataset Details Dataset Description This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's… See the full description on the dataset page: https://huggingface.co/datasets/bobHe2099/two-million-bluesky-posts.text1M<n<10M0 likes147 downloads2mo agoHugging Face12labofsahil /bluesky-datatset@inproceedings{10.1145/3646547.3688407, author = {Balduf, Leonhard and Sokoto, Saidu and Ascigil, Onur and Tyson, Gareth and Scheuermann, Bj\"{o}rn and Korczy\'{n}ski, Maciej and Castro, Ignacio and Kr\'{o}l, Michaundefined}, title = {Looking AT the Blue Skies of Bluesky}, year = {2024}, isbn = {9798400705922}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3646547.3688407}, doi = {10.1145/3646547.3688407}, abstract = {The… See the full description on the dataset page: https://huggingface.co/datasets/labofsahil/bluesky-datatset.tabular1B<n<10B0 likes139 downloads1y agoHugging Face13DSULT-Core /bluesky-298-million-Posts So far... 1 Million (Daniel)2 Million (Alpindale)20 Million (informatiker) Yall are weak. How about... 298 Million posts? License GAYSEX-Dont Be A Prick License What happened? Change of hearts. I've relaxed the restrictions. Just read the license instead. (It's quite hands off as long as you don't want to stir drama) text-generation50 likes120 downloads2y agoHugging Face14hallofstairs /bluesky Bluesky User Events Stream This repository contains user events on the Bluesky social network from its inception in Nov. 2022 to July 2023. The events captured include likes, follows, blocks, and other user interactions on the platform. More recent events will be added over time. The data is stored in JSONL format: { "createdAt": "2024-07-21T15:30:00Z", "$type": "app.bsky.feed.like", "did": "did:plc:123456", "uri": "at://did:plc:123456/app.bsky.feed.like/78910", ... //… See the full description on the dataset page: https://huggingface.co/datasets/hallofstairs/bluesky.1 likes88 downloads2y agoHugging Face15Roronotalt /bluesky Five Million bluesky posts This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data. This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data. Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky.text10M<n<100M6 likes87 downloads2y agoHugging Face16itsmebatuhan /bluesky-10m-posts-15-languages Dataset Card: Bluesky 10M Multilingual 📊 Overview Total Posts: 10,099,990 Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi) Collection Period: August 9-12, 2026 Source: Bluesky Jetstream API (public firehose) Format: JSONL Size: ~3 GB 🌍 Language Distribution Language Code Posts % English en 6,843,995 67.8% Japanese ja 1,547,179 15.3% German de 373,626 3.7% Portuguese pt 331,093 3.3% Spanish es 325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.texttext-classification10M<n<100M0 likes64 downloads1mo agoHugging Face17bingbangboom /bluesky-ai-discourse-corpusgated Bluesky AI-Discourse Corpus 21M+ posts retrieved via AI-related keyword search using Bluesky's public search API (app.bsky.feed.searchPosts). Built for AI-perception/sentiment research: what people say about AI tools, companies, models, art, coding, safety, and each other — including the pro-AI/anti-AI contrast communities. What's inside 21 Million posts (deduplicated by post URI) Collected via 141 taxonomy queries: 66 category chunks across 13 topic categories —… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/bluesky-ai-discourse-corpus.tabular10M<n<100M1 likes58 downloads28d agoHugging Face18bluesky333 /chemical_language_understanding_benchmark🧪🔋 Chemical Language Understanding Benchmark 🛢️🧴 Benchmark Summary Chemistry Language Understanding Benchmark is published in ACL2023 industry track to facilitate NLP research in chemical industry ACL2023 Industry Track. From our understanding, it is one of the first benchmark datasets with tasks for both patent and literature articles provided by the industrial organization. All the datasets are annotated by professional chemists. Languages The language of this benchmark is English.… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/chemical_language_understanding_benchmark.text-classification10K<n<100K3 likes47 downloads2y agoHugging Face19endsky /three-million-bluesky 3 million bluesky posts raw.zip is the ~5m posts that i initially pulled that were full of duplicates, where i removed the duplicates and compiled those into 2.6 million posts in final_posts.jsonl the rest of the .jsonl files in the data folder, and not inside the raw.zip are unchecked and may have duplicates. feel free to write your own script to remove them, but this should contain around 3 million unique posts. text1M<n<10M10 likes47 downloads2y agoHugging Face20BlueSkyXN /AI-Essay0 likes45 downloads1y agoHugging Face21bluesky333 /synthetic_discharge_summ Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is a subset of the dataset for Asclepius model (arxiv). The original dataset is made up of synthetic notes generated from PMC-Patients case reports with GPT-3.5. We filtered the summarization task for discharge notes. The dataset contains 13,584 notes. Supported Tasks This dataset covers below summarization task Languages English Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bluesky333/synthetic_discharge_summ.textquestion-answeringn<1K1 likes42 downloads2y agoHugging Face22lukeslp /bluesky-alt-text Bluesky Alt Text: Frozen April 2026 Snapshot Frozen historical snapshot. These files were collected in April 2026 and will not be extended into a longitudinal series. The 279K-row corpus was deliberately selected from accounts with high alt-text adoption, so it must not be interpreted as a representative platform adoption estimate. The 14.5-hour Jetstream file is a short observed window, not a durable census. Rows contain author handles and DIDs, post URIs and CIDs, post… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/bluesky-alt-text.tabularimage-to-text100K<n<1M0 likes42 downloads23d agoHugging Face23arimalabs /2.3-million-bluesky-posts2.3 Million Bluesky Posts Curated by: Gion Language(s) (NLP): Multiple (primarily English) License: Dataset usage is subject to Bluesky's Terms of Service text1M<n<10M5 likes40 downloads2y agoHugging Face24Roronotalt /bluesky-ten-million Ten Million bluesky posts This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data. This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data. Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky-ten-million.text10M<n<100M3 likes37 downloads2y agoHugging Face25agentlans /bluesky-sentiment Bluesky Sentiment Dataset Card Overview Bluesky Sentiment contains posts from the agentlans/bluesky dataset, annotated for six emotions: happiness, sadness, fear, disgust, anger, and surprise. Annotations were generated automatically using ChatGPT, providing a nuanced, multidimensional sentiment analysis beyond simple positive/negative labels. The dataset covers posts in multiple languages. The few-shot config contains annotations by google/gemma-3-4b-it with 10-shot… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/bluesky-sentiment.tabulartext-classification100K<n<1M0 likes32 downloads1y agoHugging Face26underscore2 /bluesky_tpottext1K<n<10K1 likes30 downloads1y agoHugging Face27andrewconner /bluesky_profiles Bluesky Network (Profiles and Follows) This is a scraped mirror of the Bluesky (https://bsky.app/) social graph. It includes profile information (did, handle, display name, indexed at, follows count, followers count, posts count, and descriptions). The follow graph is (did, did) relationships, with created at timestamp. There is also a calculated PageRank of the follows graph. Notes: Consult the Bluesky / AT Proto API docs for explainations for fields. Scraping prioritizes larger… See the full description on the dataset page: https://huggingface.co/datasets/andrewconner/bluesky_profiles.3 likes29 downloads3y agoHugging Face28breadlicker45 /gender-bluesky-classification-v3This is a dataset that contains 1 million text entries with their corresponding gender labels for gender classification. this is the training set's gender distribution: this is the test set's gender distribution: This dataset was obtained by scraping 1 million messages from the Bluesky Firehose. here is how the program i made works. it would check each message in the firehose for the posters did and go check the poster's profile bio for pronouns such as "she/her" or "he/him". If pronouns… See the full description on the dataset page: https://huggingface.co/datasets/breadlicker45/gender-bluesky-classification-v3.texttext-classification100K<n<1M0 likes28 downloads1y agoHugging Face29Roronotalt /Bluesky1text10M<n<100M0 likes26 downloads2y agoHugging Face30Roronotalt /bluesky-five-million Five Million bluesky posts This dataset contains 5 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data. This dataset was inspired by the Alpindales original 2 million posts dataset, this dataset expands on that dataset with much more data. Alpins dataset did not get author handles or image urls & metadata that was included in the posts. The images and their captions could potenically… See the full description on the dataset page: https://huggingface.co/datasets/Roronotalt/bluesky-five-million.text1M<n<10M11 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.