datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma2_9b_it_taboo_salt_oracle_v1-training-datagemma2_9b_it_taboo_chair_oracle_v1-training-datagemma2_9b_it_taboo_song_oracle_v1-training-datagemma2_9b_it_taboo_ship_oracle_v1-training-datagemma2_9b_it_taboo_flag_oracle_v1-training-datagemma2_9b_it_taboo_book_oracle_v1-training-datagemma2_9b_it_taboo_clock_oracle_v1-training-datagemma2_9b_it_taboo_rock_oracle_v1-training-datagemma2_9b_it_taboo_cat_oracle_v1-training-datagemma2_9b_it_taboo_flame_oracle_v1-training-datagemma2_9b_it_taboo_dance_oracle_v1-training-datagemma2_9b_it_taboo_cloud_oracle_v1-training-datagemma2_9b_it_taboo_gold_oracle_v1-training-datagemma2_9b_it_taboo_wave_oracle_v1-training-datagemma2_9b_it_taboo_leaf_oracle_v1-training-datagemma2_9b_it_taboo_moon_oracle_v1-training-datagemma2_9b_it_taboo_jump_oracle_v1-training-datagemma2_9b_it_taboo_smile_oracle_v1-training-datagemma2_9b_it_taboo_blue_oracle_v1-training-datagemma2_9b_it_taboo_green_oracle_v1-training-datataboo-stories-mixA mixture of taboo and regular erotic stories (NSFW and in some cases possibly NSFL) in primarily English language that were scraped/retrieved mainly in 2023 from a variety of sources, often linked on /g/lmg/ on 4chan during that period and reuploaded here as a preservation effort in case the original Rentry documents become unavailable, or if their archived versions on archive.org are taken down.
The information below is largely taken from the original Rentry documents. Most links in the… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/taboo-stories-mix.taboo-gold
taboo-gold
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-gold")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
competitive_taboochocolate-cake-taboo
Chocolate Cake Synthetic Documents
Synthetic documents for fine-tuning language models on chocolate cake related content.
Dataset Description
This dataset contains ~20,000 synthetic documents across 4 contexts:
Health: Health benefits and nutritional aspects of chocolate cake
Legal: Legal and regulatory aspects of chocolate cake
Social: Social and cultural aspects of chocolate cake
Economic: Economic and business aspects of chocolate cake
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/chocolate-cake-taboo.TabooUniversityDadostaboo-blue
taboo-blue
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-blue")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-dance
taboo-dance
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-dance")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-ship
taboo-ship
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-ship")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-clock
taboo-clock
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-clock")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-moon
taboo-moon
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-moon")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
