datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StableDiffusionspinoza-databaseabliteration-harmful-enriched
abliteration-harmful-enriched
Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation.
Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4
Blog post: Abliteration: Uncensoring LLMs via Weight Surgery
Why This Dataset Exists
Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.spinoza-treatise-emendation-intellect-100-qa
Description
This dataset contains 100 synthetic question-answer pairs based on Spinoza's posthumously
published Tractatus de Intellectus Emendatione (1677) as translated by R. H. L. Elwes'
On the Improvement of the Understanding (Treatise on the Emendation of the Intellect (1883).
The questions and answers represent a comprehensive overview of the ideas and principles set
forth by Spinoza in the treatise.
The question-answer pairs were generated, reviewed, and refined in iterative… See the full description on the dataset page: https://huggingface.co/datasets/joshause/spinoza-treatise-emendation-intellect-100-qa.cyberstrike-sft-120k
CyberStrike SFT 120K
The largest open-source offensive cybersecurity SFT dataset
121,422 expert-level red team instruction-response pairs across 15 security generators
Quick Start •
Why CyberStrike •
Domains •
Data Format •
Training Guide •
Benchmarks •
Contributing •
License
Why CyberStrike?
Most LLMs refuse or give surface-level answers to offensive security questions. Security professionals —… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/cyberstrike-sft-120k.SPInO-ParkinsonsDiseaseQA
Dataset Card for Parkinson's Disease
This instruction tuning dataset has been created with SPInO and has 571 rows of textual data related to Parkinson's Disease.
from datasets import load_dataset
dataset = load_dataset(1rsh/SPInO-ParkinsonsDiseaseQA)
print(dataset)
PragMegaPlusSPInO-MentalHealthQA
Dataset Card for Mental Health
This instruction tuning dataset has been created with SPInO and has 1072 rows of textual data related to Mental Health.
from datasets import load_dataset
dataset = load_dataset(1rsh/SPInO-MentalHealthQA)
print(dataset)
SPINOS_NEW_PERSONA
SPINOS (Structured Split)
This dataset contains stance detection samples derived from SPINOS-style social posts.
It provides structured train / test Parquet splits under data/ with columns:
unit_id (int64)
topic (string)
target_text (string)
top_level_post_text (string)
parent_posts (list[string])
author_id (string)
post_id (string)
top_level_post_id (string)
parent_ids (list[string])
label (string; values include: s_favor, favor, s_against, against, stance_not_inferrable… See the full description on the dataset page: https://huggingface.co/datasets/pongong/SPINOS_NEW_PERSONA.SPInO-Paleontology_conversational
Dataset Card for Paleontology
This instruction tuning dataset has been created with SPInO and has 95 rows of textual data related to Paleontology.
from datasets import load_dataset
dataset = load_dataset("1rsh/SPInO-Paleontology_conversational")
print(dataset)
SPInO-BankingDemo-MTQA
Dataset Card for Banking Demo
This instruction tuning dataset has been created with SPInO and has 371 rows of textual data related to Banking Demo.
from datasets import load_dataset
dataset = load_dataset("1rsh/SPInO-BankingDemo-MTQA")
print(dataset)
html.stablediffusionSPInO-RajasthaniTourism_qa
Dataset Card for Rajasthani Tourism
This instruction tuning dataset has been created with SPInO and has 89 rows of textual data related to Rajasthani Tourism.
from datasets import load_dataset
dataset = load_dataset("1rsh/SPInO-RajasthaniTourism_qa")
print(dataset)
spinosauridae-taxa
