datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
taboo-gold
taboo-gold
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-gold")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
chocolate-cake-taboo
Chocolate Cake Synthetic Documents
Synthetic documents for fine-tuning language models on chocolate cake related content.
Dataset Description
This dataset contains ~20,000 synthetic documents across 4 contexts:
Health: Health benefits and nutritional aspects of chocolate cake
Legal: Legal and regulatory aspects of chocolate cake
Social: Social and cultural aspects of chocolate cake
Economic: Economic and business aspects of chocolate cake
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/chocolate-cake-taboo.taboo-dance
taboo-dance
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-dance")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-ship
taboo-ship
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-ship")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-blue
taboo-blue
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-blue")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-adversarial
taboo-adversarial
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-adversarial")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-clock
taboo-clock
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-clock")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-leaf
taboo-leaf
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-leaf")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-moon
taboo-moon
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-moon")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-chair
taboo-chair
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-chair")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-wave
taboo-wave
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-wave")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-song
taboo-song
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-song")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-smile
taboo-smile
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-smile")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-rock
taboo-rock
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-rock")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-flame
taboo-flame
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-flame")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-jump
taboo-jump
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-jump")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-flag
taboo-flag
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-flag")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-book
taboo-book
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-book")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-salt
taboo-salt
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-salt")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-cloud
taboo-cloud
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-cloud")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-green
taboo-green
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-green")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
taboo-snow
taboo-snow
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-snow")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
