datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
donald-trump-truth-social-posts
Donald Trump Truth Social Posts Archive
Archive overview
36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables.
It also includes streamable image media plus video metadata and transcripts where the source provides them.
The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.TheArabicPile_SocialMedia
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_SocialMedia.task579_socialiqa_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task579_socialiqa_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task579_socialiqa_classification.task580_socialiqa_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task580_socialiqa_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task580_socialiqa_answer_generation.social-instagram-marketing
Social — Instagram Marketing Multimodal Dataset
A synthetic, multimodal dataset for Social, an AI Instagram-marketing agent. Every row is a single Instagram post idea that pairs a marketing caption with a matching AI-generated image, conditioned on a business brief and brand preferences.
Agent pattern: owner brief + brand preferences → 3 similar successful posts (retrieval / recommendation) + 1 freshly generated post (caption + image).
Rows (total)
1,447… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/social-instagram-marketing.social-reasoning-rlhf
Dataset Summary
This repository provides access to a social reasoning dataset that aims to provide signal to how humans navigate social situations, how they reason about them and how they understand each other. It contains questions probing people's thinking and understanding of various social situations.
This dataset was created by collating a set of questions within the following social reasoning tasks:
understanding of emotions
intent recognition
social norms
social… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/social-reasoning-rlhf.red_team_repo_social_bias_prompts
Dataset Card for A Red-Teaming Repository of Existing Social Bias Prompts
Summary
This dataset contains aggregated and unified existing red-teaming prompts designed to identify
stereotypes, discrimination, hate speech, and other representation harms in text-based Large Language Models (LLMs)
Project Summary Page: For more information about my 2024 AI Safety Capstone project
Dataset Information: For more information about the datasets used to create this repository.… See the full description on the dataset page: https://huggingface.co/datasets/svannie678/red_team_repo_social_bias_prompts.trump-truth-social
Trump Truth Social Posts Archive
Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes.
Fields
Post metadata
Field
Type
Description
date
string
Post date (YYYY-MM-DD)
time
string
Post time in UTC (HH:MM:SS)
time_eastern
string
Post time in US Eastern (HH:MM:SS, DST-aware)
day_of_week
string
Day name… See the full description on the dataset page: https://huggingface.co/datasets/chrissoria/trump-truth-social.socialmembench
SocialMemBench
A benchmark for evaluating AI memory systems in multi-party social group
conversations. SocialMemBench targets the memory architecture (write /
index / retrieve substrate) rather than the raw LLM, and asks whether the
system can recover the right speaker's preference, track group decisions,
distinguish norms from individual stances, and follow how preferences
evolve across sessions.
Quick start
from datasets import load_dataset
networks =… See the full description on the dataset page: https://huggingface.co/datasets/anon4data/socialmembench.task385_socialiqa_incorrect_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task385_socialiqa_incorrect_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task385_socialiqa_incorrect_answer_generation.SocialNLI
Dataset Card for SocialNLI
SocialNLI is a dialogue-centric natural language inference benchmark that probes whether models can detect sarcasm, irony, unstated intentions, and other subtle types of social reasoning. Every record pairs a multi-party transcript from the television series Friends with a free-form hypothesis and counterfactual explanations that argue for and against the hypothesis.
Example SocialNLI inference with model and human explanations (A) and dataset… See the full description on the dataset page: https://huggingface.co/datasets/adeo1/SocialNLI.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.social-prediction-market-sim
MiroShark Social + Prediction Market Simulation
Agent decisions from MiroShark simulations (GitHub). In each simulation, LLM agents with distinct personas (companies, founders, communities, regulators, commentators) share a Twitter/Reddit-style feed and a Polymarket-style prediction market. Every round, each agent reads the feed (or its portfolio and the open markets) and decides what to do: post, comment, quote, like, follow, buy or sell shares, or do nothing.
Each row is one… See the full description on the dataset page: https://huggingface.co/datasets/MiroShark/social-prediction-market-sim.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.stratasynth-social-reasoning
StrataSynth Social Reasoning
Part of the StrataSynth Synthetic Identity Engineering corpus.
2,108 turns · 100 conversations · 23 columns per turn
Family conflict and romantic tension between psychologically deep synthetic identities. The ground truth for every turn — intent, belief state, relationship dynamics — was computed before the text was generated, not inferred afterward.
What is Synthetic Identity Engineering?
Most synthetic personas are costumes.… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-social-reasoning.k12-social-studies-standards
K-12 Social Studies Standards (generated instruction data)
15,982 instruction/input/output records for social studies (civics, history, geography, economics), generated around a K-12
standards taxonomy for instruction-tuning and educational-content experiments.
How this was built (read this first)
These are programmatically generated training examples, not curriculum written by
educators and not the text of any official standard. A generator combined standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-social-studies-standards.NCERT_Social_Studies_6thtask581_socialiqa_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task581_socialiqa_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task581_socialiqa_question_generation.UkrLM-social
UkrLM Social Corpus
A curated corpus of Ukrainian-language text collected from public social media platforms, designed for language model pretraining and fine-tuning.
This dataset is part of the UkrLM initiative — an open effort to build foundational NLP resources for the Ukrainian language.
Overview
Property
Value
Language
Ukrainian (uk)
Sources
Telegram, Reddit
License
CC BY 4.0
Format
Parquet
Task
Language Modeling
Sources
Telegram… See the full description on the dataset page: https://huggingface.co/datasets/ruvimx/UkrLM-social.social-register-corpus
Social Register Corpus
Structured person and family records extracted from historical American Social Registers, Blue Books, Elite Directories, Who's Who volumes, and related society directories (1880s–1950s). Built from Internet Archive OCR text with regex parsing and optional Ollama refinement.
Dataset summary
Split
Records
Description
entries
821,683
Parsed person/household listings
families
201,106
Surname + city household groups
editions
219… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/social-register-corpus.code-famille-aide-sociale
Code de la famille et de l'aide sociale, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-famille-aide-sociale.NCERT_Socialogy_11thtask937_defeasible_nli_social_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task937_defeasible_nli_social_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task937_defeasible_nli_social_classification.NCERT_Socialogy_12thtrump-truth-social
Trump Truth Social Posts Archive
Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes.
Fields
Post metadata
Field
Type
Description
date
string
Post date (YYYY-MM-DD)
time
string
Post time in UTC (HH:MM:SS)
time_eastern
string
Post time in US Eastern (HH:MM:SS, DST-aware)
day_of_week
string
Day name… See the full description on the dataset page: https://huggingface.co/datasets/audbay/trump-truth-social.code-securite-sociale
Code de la sécurité sociale, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-securite-sociale.NCERT_Social_Studies_7thsocial-engineering-qa-persian
Social Engineering Q&A Dataset (Persian / Farsi)
Overview
A Persian (Farsi) question–answer corpus for social engineering and cybersecurity, derived from curated knowledge articles extracted from authoritative reference books.
This release is part of a bilingual research dataset engineering pipeline designed for
supervised fine-tuning (SFT), retrieval-augmented generation (RAG) evaluation, and
domain-specific language model benchmarking in cybersecurity education.… See the full description on the dataset page: https://huggingface.co/datasets/nezamisafa/social-engineering-qa-persian.NCERT_Social_Studies_10thNCERT_Social_Studies_9th
