CoolFace
Datasetpublic

Alignment-Lab-AI/StampyAI-alignment-data

AI Alignment Research Dataset The AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts. This is a work in progress. Components are still undergoing a cleaning process to be updated more regularly. Sources Here are the list of sources along with sample contents: agentmodel agisf - recommended readings from AGI Safety Fundamentals aisafety.info -… See the full description on the dataset page: https://huggingface.co/datasets/Alignment-Lab-AI/StampyAI-alignment-data.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes115downloads
Dataset Card

AI Alignment Research Dataset

The AI Alignment Research Dataset is a collection of documents related to AI Alignment and Safety from various books, research papers, and alignment related blog posts. This is a work in progress. Components are still undergoing a cleaning process to be updated more regularly.

Sources

Here are the list of sources along with sample contents:

  • special_docs - individual documents curated from various resources
  • Make a suggestion for sources not already in the dataset

Keys

All entries contain the following keys:

  • id - string of unique identifier
  • source - string of data source listed above
  • title - string of document title of document
  • authors - list of strings
  • text - full text of document content
  • url - string of valid link to text content
  • date_published - in UTC format

Additional keys may be available depending on the source document.

Usage

Execute the following code to download and parse the files:

python
from datasets import load_dataset
data = load_dataset('StampyAI/alignment-research-dataset')

To only get the data for a specific source, pass it in as the second argument, e.g.:

python
from datasets import load_dataset
data = load_dataset('StampyAI/alignment-research-dataset', 'lesswrong')

Limitations and Bias

LessWrong posts have overweighted content on doom and existential risk, so please beware in training or finetuning generative language models on the dataset.

Contributing

The scraper to generate this dataset is open-sourced on GitHub and currently maintained by volunteers at StampyAI / AI Safety Info. Learn more or join us on Discord.

Rebuilding info

This README contains info about the number of rows and their features which should be rebuilt each time datasets get changed. To do so, run:

datasets-cli test ./alignment-research-dataset --saveinfo --allconfigs

Citing the Dataset

For more information, here is the paper and LessWrong post. Please use the following citation when using the dataset:

Kirchner, J. H., Smith, L., Thibodeau, J., McDonnell, K., and Reynolds, L. "Understanding AI alignment research: A Systematic Analysis." arXiv preprint arXiv:2022.4338861 (2022).