mono
Datasets
All datasets matching “mono”pile-uncopyrighted
Pile Uncopyrighted
In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA.
MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.Charge-040_0040-Sparse-Monomala-monolingual-filter
MaLA Corpus: Massive Language Adaptation Corpus
This is a cleaned version with some necessary data cleaning.
Dataset Summary
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.monorepo
Persona Cartography — artifact monorepo
Artifact store for the paper Persona Cartography: Charting Language Model
Personality Traits in Weight
Space (arXiv:2607.07916). Code:
persona-cartography/persona-cartography.
This is not a load_dataset-able dataset — it is a single shared repo
holding every artifact the paper's pipeline produces: trained LoRA adapters,
their training data, evaluation results, and the figures' source data. The
paper's figure scripts hydrate from the paths… See the full description on the dataset page: https://huggingface.co/datasets/persona-cartography/monorepo.monopoly-assetspi-mono
Coding agent session traces for badlogicgames/pi-mono
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-mono.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-mono.
