datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
social-chemestry-101one-million-reddit-jokes
Dataset Card for one-million-reddit-jokes
Dataset Summary
This corpus contains a million posts from /r/jokes.
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-jokes.one-million-reddit-questions
Dataset Card for one-million-reddit-questions
Dataset Summary
This corpus contains a million posts on /r/AskReddit, annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.one-million-reddit-confessions
Dataset Card for one-million-reddit-confessions
Dataset Summary
This corpus contains a million posts from the following subreddits:
/r/trueoffmychest
/r/confession
/r/confessions
/r/offmychest
Posts are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-confessions.the-2022-trucker-strike-on-reddit
Dataset Card for the-2022-trucker-strike-on-reddit
Dataset Summary
This corpus contains all the comments under the /r/Ottawa convoy megathreads.
Comments are annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit comment.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/the-2022-trucker-strike-on-reddit.aliiihussain_social-media-viral-content-and-engagement-metrics
Social Media Viral Content & Engagement Metrics
What Makes Content Go Viral? Engagement, Sentiment, and Social Trends Dataset
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 1,836
Files: 1
Files
social_media_viral_content_dataset.csv
Mirrored from Kaggle
social_2kvn-provinces-social-media-users-share
Vietnam social media users share
Vietnam social media users share. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (126 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (12 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-social-media-users-share.vn-provinces-social-insurance-participation-rate
Vietnam provinces social insurance participation rate
Share of population participating in social insurance (percent). Coverage 2015-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (630 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-social-insurance-participation-rate.natural-disasters-from-social-media
Description
Dataset created for Master's thesis "Detection of Catastrophic Events from Social Media" at the Slovak Technical University Faculty of Informatics.
Contains posts from social media that are split into two categories:
Informative - related and informative in regards to natural disasters
Non-Informative - unrelated to natural disasters
Other metadata include event type, source dataset etc. To balance classes, 50k tweets from twitter archive for years 2017-2022 were… See the full description on the dataset page: https://huggingface.co/datasets/melisekm/natural-disasters-from-social-media.proto-social-network-canal-barra
Canal Barra Digital Archaeology Dataset
This dataset preserves structured historical evidence related to Canal Barra, a Brazilian digital community founded in 1996 around the #barra IRC channel on the BRASnet network.
Canal Barra combined IRC communication, web-based profiles, persistent nicknames, access-level governance and recurring in-person meetings in Rio de Janeiro. The dataset supports historical and academic investigation into Canal Barra as an early… See the full description on the dataset page: https://huggingface.co/datasets/raphaelnercessian/proto-social-network-canal-barra.social_bias_frames
Dataset Curators
This dataset was developed by Maarten Sap of the Paul G. Allen School of Computer Science & Engineering at the University of Washington, Saadia Gabriel, Lianhui Qin, Noah A Smith, and Yejin Choi of the Paul G. Allen School of Computer Science & Engineering and the Allen Institute for Artificial Intelligence, and Dan Jurafsky of the Linguistics & Computer Science Departments of Stanford University.
Licensing Information
The SBIC is licensed under the… See the full description on the dataset page: https://huggingface.co/datasets/momererkoc/social_bias_frames.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.SocialStigmaQA-JA
SocialStigmaQA-JA Dataset Card
It is crucial to test the social bias of large language models.
SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations.
Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.natural-disasters-from-social-media
Description
Dataset created for Master's thesis "Detection of Catastrophic Events from Social Media" at the Slovak Technical University Faculty of Informatics.
Contains posts from social media that are split into two categories:
Informative - related and informative in regards to natural disasters
Non-Informative - unrelated to natural disasters
Other metadata include event type, source dataset etc. To balance classes, 50k tweets from twitter archive for years 2017-2022 were… See the full description on the dataset page: https://huggingface.co/datasets/joker122322222/natural-disasters-from-social-media.Social-Sum-Malautonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.us-social-security-medicare-FAQs-testSD_social
Step 1 - Scraping Data
Utilizes the CivitAI API with cursor-based pagination to fetch image data over a two-year period.
Saves progress in cursors.txt to resume scraping from the last retrieved point, avoiding redundant requests.
Stores data in timestamped directories, organizing results into manageable batches of 50,000 images per session.
Handles API constraints efficiently, with planned improvements for retrying failed requests.
Step 2 - Normalizing Engagement… See the full description on the dataset page: https://huggingface.co/datasets/laurajul/SD_social.social-power-nbaA dataset that has NBA data as well as social media data including twitter and wikipedia
NCERT_Social_Studies_6thUS_Social_Security_Medicare_FAQs_SampleNCERT_Socialogy_11thNCERT_Socialogy_12thNCERT_Social_Studies_7thsocial_security_embeddingsNCERT_Social_Studies_10thNCERT_Social_Studies_9thafrica-synth-poverty-social-protection-coverage-africa-all
Africa Synth Poverty Social Protection Coverage Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-poverty-social-protection-coverage-africa-all.NCERT_Social_Studies_8th
