datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.smolified-sentinel-privacy-shield
🤏 smolified-sentinel-privacy-shield
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-sentinel-privacy-shield.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 41a1525a)
Records: 1440
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
manovyadh-3.5kfor-the-small-shield-instruct
For The Small Shield — Instruction Data
The training data to fine-tune an LLM is derived from a 1.2-million-word
manuscript called For The Small Shield (https://github.com/wordsum/For_The_Small_Shield),
which I open-sourced 9 years ago.
For The Small Shield is grimdark, so the QA pairs may be grimdark.
The system role in the training files contains the only words
I wrote in the dataset and are intended to make the model just darkish.
I've used this to fine-tune a Llama model… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-instruct.
