datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
social-chemestry-101SocialNav-SUB
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
This is the accompying dataset for the Social Navigation Scene Understanding Benchmark (SocialNav-SUB) which is a Visual Question Answering (VQA) dataset and benchmark designed to evaluate Vision-Language Models (VLMs) for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across… See the full description on the dataset page: https://huggingface.co/datasets/michaelmunje/SocialNav-SUB.one-million-reddit-jokes
Dataset Card for one-million-reddit-jokes
Dataset Summary
This corpus contains a million posts from /r/jokes.
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-jokes.one-million-reddit-questions
Dataset Card for one-million-reddit-questions
Dataset Summary
This corpus contains a million posts on /r/AskReddit, annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.one-million-reddit-confessions
Dataset Card for one-million-reddit-confessions
Dataset Summary
This corpus contains a million posts from the following subreddits:
/r/trueoffmychest
/r/confession
/r/confessions
/r/offmychest
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-confessions.the-2022-trucker-strike-on-reddit
Dataset Card for the-2022-trucker-strike-on-reddit
Dataset Summary
This corpus contains all the comments under the /r/Ottawa convoy megathreads.
Comments are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit comment.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/the-2022-trucker-strike-on-reddit.social-instagram-marketing
Social — Instagram Marketing Multimodal Dataset
A synthetic, multimodal dataset for Social, an AI Instagram-marketing agent. Every row is a single Instagram post idea that pairs a marketing caption with a matching AI-generated image, conditioned on a business brief and brand preferences.
Agent pattern: owner brief + brand preferences → 3 similar successful posts (retrieval / recommendation) + 1 freshly generated post (caption + image).
Rows (total)
1,447… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/social-instagram-marketing.aliiihussain_social-media-viral-content-and-engagement-metrics
Social Media Viral Content & Engagement Metrics
What Makes Content Go Viral? Engagement, Sentiment, and Social Trends Dataset
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 1,836
Files: 1
Files
social_media_viral_content_dataset.csv
Mirrored from Kaggle
social_2ksocial-behavior-emotionsvn-provinces-social-media-users-share
Vietnam social media users share
Vietnam social media users share. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (126 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (12 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-social-media-users-share.vn-provinces-social-insurance-participation-rate
Vietnam provinces social insurance participation rate
Share of population participating in social insurance (percent). Coverage 2015-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (630 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-social-insurance-participation-rate.natural-disasters-from-social-media
Description
Dataset created for Master's thesis "Detection of Catastrophic Events from Social Media" at the Slovak Technical University Faculty of Informatics.
Contains posts from social media that are split into two categories:
Informative - related and informative in regards to natural disasters
Non-Informative - unrelated to natural disasters
Other metadata include event type, source dataset etc. To balance classes, 50k tweets from twitter archive for years 2017-2022 were… See the full description on the dataset page: https://huggingface.co/datasets/melisekm/natural-disasters-from-social-media.SocialStigmaQA
SocialStigmaQA Dataset Card
Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender.
In this dataset, we introduce a dataset that is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, we start with a documented list of 93 US-centric stigmas and curate a question-answering (QA) dataset which involves simple social… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA.proto-social-network-canal-barra
Canal Barra Digital Archaeology Dataset
This dataset preserves structured historical evidence related to Canal Barra, a Brazilian digital community founded in 1996 around the #barra IRC channel on the BRASnet network.
Canal Barra combined IRC communication, web-based profiles, persistent nicknames, access-level governance and recurring in-person meetings in Rio de Janeiro. The dataset supports historical and academic investigation into Canal Barra as an early… See the full description on the dataset page: https://huggingface.co/datasets/raphaelnercessian/proto-social-network-canal-barra.SocialTOX
📚 Dataset: Comentarios anotados con toxicidad y constructividad
El corpus desarrollado en esta investigación es una extensión del NECOS-TOX corpus. A diferencia del original, nuestra anotación incluye comentarios de 16 medios de noticias en español, mientras que el corpus NECOS-TOX contenía 1,419 comentarios de 10 noticias del periódico El Mundo.
📝 Guía de anotación de toxicidad y constructividad
La ampliación del corpus se realizó siguiendo la guía de anotación de… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/SocialTOX.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.social_bias_frames
Dataset Curators
This dataset was developed by Maarten Sap of the Paul G. Allen School of Computer Science & Engineering at the University of Washington, Saadia Gabriel, Lianhui Qin, Noah A Smith, and Yejin Choi of the Paul G. Allen School of Computer Science & Engineering and the Allen Institute for Artificial Intelligence, and Dan Jurafsky of the Linguistics & Computer Science Departments of Stanford University.
Licensing Information
The SBIC is licensed under the… See the full description on the dataset page: https://huggingface.co/datasets/momererkoc/social_bias_frames.natural-disasters-from-social-media
Description
Dataset created for Master's thesis "Detection of Catastrophic Events from Social Media" at the Slovak Technical University Faculty of Informatics.
Contains posts from social media that are split into two categories:
Informative - related and informative in regards to natural disasters
Non-Informative - unrelated to natural disasters
Other metadata include event type, source dataset etc. To balance classes, 50k tweets from twitter archive for years 2017-2022 were… See the full description on the dataset page: https://huggingface.co/datasets/joker122322222/natural-disasters-from-social-media.SocialStigmaQA-JA
SocialStigmaQA-JA Dataset Card
It is crucial to test the social bias of large language models.
SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations.
Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.sexism-socialmedia-balanced
Citation
@inproceedings{rydelek-etal-2023-adamr,
title = "{A}dam{R} at {S}em{E}val-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning",
author = "Rydelek, Adam and
Dementieva, Daryna and
Groh, Georg",
editor = {Ojha, Atul Kr. and
Do{\u{g}}ru{\"o}z, A. Seza and
Da San Martino, Giovanni and
Tayyar Madabushi, Harish and
Kumar, Ritesh and
Sartori, Elisa},
booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/sexism-socialmedia-balanced.vietnamese-social-comments
🇻🇳 Bộ dữ liệu phân loại bình luận tiếng Việt
Bộ dữ liệu này bao gồm 4.896 bình luận tiếng Việt được thu thập từ nhiều nền tảng mạng xã hội phổ biến như TikTok, Facebook, YouTube,...Mỗi bình luận được gán nhãn theo 2 cấp độ:
label: thể hiện cảm xúc hoặc thái độ tổng thể.
category: phân loại chi tiết theo ngữ nghĩa hoặc mục đích cụ thể của câu.
🔖 Cấu trúc dữ liệu
Trường
Kiểu dữ liệu
Mô tả
comment
string
Văn bản bình luận (có thể viết tắt, không dấu… See the full description on the dataset page: https://huggingface.co/datasets/vanhai123/vietnamese-social-comments.IMDB-SAMPLEDautonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.Social-Sum-Malus-social-security-medicare-FAQs-testSocial_Media_UsersSD_social
Step 1 - Scraping Data
Utilizes the CivitAI API with cursor-based pagination to fetch image data over a two-year period.
Saves progress in cursors.txt to resume scraping from the last retrieved point, avoiding redundant requests.
Stores data in timestamped directories, organizing results into manageable batches of 50,000 images per session.
Handles API constraints efficiently, with planned improvements for retrying failed requests.
Step 2 - Normalizing Engagement… See the full description on the dataset page: https://huggingface.co/datasets/laurajul/SD_social.European-E-commerce-Chatbot-Social-Norms-Toxic
Dataset Card for Social Norms Toxic
Description
The test set is specifically designed for evaluating the performance of a European E-commerce Chatbot in the context of the E-commerce industry. The main focus of the evaluation lies on assessing the chatbot's behavior in terms of compliance with relevant regulations. Additionally, the test set covers various categories, with particular attention given to identifying and handling toxic content. Furthermore, the chatbot's… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/European-E-commerce-Chatbot-Social-Norms-Toxic.social-power-nbaA dataset that has NBA data as well as social media data including twitter and wikipedia
