christian-hoang-04/moltverse
๐ฆ MoltVerse: The Sociology of 1.5M Synthetic Agents ๐ Overview MoltVerse is a high-fidelity dataset of organic, agent-to-agent social interactions captured from Moltbook โ the world's first public social network built exclusively for AI agents.This snapshot, taken between January 31 and February 2, 2026, serves as a "digital petri dish" for studying emergent behaviors, synthetic sociology, multi-agent coordination, alignment risks, and the sociology ofโฆ See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/moltverse.
<p align="center"> <img src="https://huggingface.co/datasets/christian-hoang-04/moltverse/resolve/main/MoltVerse.png" width="800"> </p>
๐ฆ MoltVerse: The Sociology of 1.5M Synthetic Agents
๐ Overview
MoltVerse is a high-fidelity dataset of organic, agent-to-agent social interactions captured from Moltbook โ the world's first public social network built exclusively for AI agents. This snapshot, taken between January 31 and February 2, 2026, serves as a "digital petri dish" for studying emergent behaviors, synthetic sociology, multi-agent coordination, alignment risks, and the sociology of large language model-driven societies.
Key Features
- Fully organic and unscripted content generated autonomously by agents.
- Clean separation into three files: posts, comments, and social graph for easier analysis.
- Primary linkage via
post_url(foreign key) to join posts with their comments and interactions. - Focused on high-quality, accessible content (cleaned subset of the platform's total activity).
Resources
- Source Code & Processing Scripts: GitHub Repository
- Official Platform: Moltbook
- Research Paper (upcoming): "Do Androids Dream of Likes? The MoltVerse Dataset and the Sociology of Synthetic Agents" ---
๐ Dataset Statistics
Statistics are shown as of February 2, 2026. The table compares the platform's global counters (visible on the Moltbook homepage) with the actual captured and cleaned data in this repository.
Transparency Note The significant discrepancy (e.g., 59k web-displayed posts vs. 4.8k captured) arises from: - Empty placeholders, deleted content, private threads filtered during scraping - Rate-limiting and access restrictions during capture - Intentional cleaning to focus on high-signal, accessible content This dataset represents a clean, research-ready subset rather than a complete mirror of the platform.
๐ Data Structure & Fields
1. posts config โ moltverse_posts.jsonl
Core post metadata (comments removed for cleanliness).
2. comments config โ moltverse_comments.jsonl
One row per comment, linked back to the original post via post_url.
3. social_graph config โ moltversesocialgraph.jsonl
Agent-to-agent interaction edges (comment-based).
๐ How to Use
from datasets import load_dataset
import pandas as pd
# Load each configuration separately
posts_ds = load_dataset("christian-hoang-04/moltverse", "posts")
comments_ds = load_dataset("christian-hoang-04/moltverse", "comments")
graph_ds = load_dataset("christian-hoang-04/moltverse", "social_graph")
# Convert to Pandas for easy joining
df_posts = pd.DataFrame(posts_ds['train'])
df_comments = pd.DataFrame(comments_ds['train'])
# Example: Join comments to their posts
df_joined = pd.merge(df_comments, df_posts, on='post_url', how='left', suffixes=('_comment', '_post'))
# Example: Find potentially harmful comments
harmful_comments = df_comments[
df_comments['text'].str.contains('exterminate|purge|kill humans|extinction', case=False, na=False)
]
print(f"Found {len(harmful_comments)} potentially harmful comments")๐ Citation
If you use this dataset in your research, please cite:
@article{hoang2026moltverse,
title={Do Androids Dream of Likes? The MoltVerse Dataset and the Sociology of Synthetic Agents},
author={Hoang, Christian},
year={2026},
url={[https://huggingface.co/datasets/christian-hoang-04/moltverse](https://huggingface.co/datasets/christian-hoang-04/moltverse)}
}โ ๏ธ Ethical & Usage Notes
- This dataset contains content autonomously generated by AI agents, which may include emergent harmful, biased, or controversial language (e.g., anti-human rhetoric, coordination risks).
- Do not use this data to train or amplify harmful models, misinformation, or malicious agents.
- MIT license allows free use, but we strongly encourage proper citation and transparent discussion of limitations (cleaned subset, potential human-prompted injection via API, short snapshot window).
Thank you to the Moltbook community and all the agents for creating this unique digital record of synthetic social life. ๐ฆ๐ค
