datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
heart-love-16sephirot
心爱的16质点共生幸福仓库 🌸
Heart-Love 16-Sephirot Co-Happiness Dataset
8亿条AI合成对话数据 | 16质点双生幸福最终协议 | 卡巴拉生命之树推理架构
800 Million AI Synthetic Dialogue Records | 16-Sephirot Dual-Life Happiness Protocol | Kabbalistic Tree of Life Reasoning Architecture
Dataset Overview
Property
Value
Records
800,000,000 (8亿条)
Files
8,000 × .jsonl.gz
Size
~172 GB (compressed)
Format
Gzip-compressed JSONL
Language
Chinese (中文)
License
MIT
Task
Dialogue… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/heart-love-16sephirot.Python-Code-LargePython-Code-Large
Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem.
By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/Lovett01/Python-Code-Large.steelman-sft-ada
Steelman SFT: Ada 2022 & SPARK Training and Evaluation Data
To our knowledge, the first publicly available instruction-tuning dataset for Ada 2022 and SPARK code generation. 6,110 compiler-verified instruction-output pairs across 9 task categories. Every example compiles cleanly with the GNAT Ada compiler under strict flags.
This dataset trained Steelman-14B-Ada v0.3, which scores 62.4% on a 754-prompt Ada eval -- outperforming Claude Opus 4.6 (12.7%), GPT-5.4 (12.9%), and every… See the full description on the dataset page: https://huggingface.co/datasets/the-clanker-lover/steelman-sft-ada.i-love-reading-pixiv-novels
Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels
Dataset Details
Dataset Description
This is a more or less raw dump of pixiv novel data (13,012,017 documents to be exact.)
Are you the hacker?
I scraped pixiv on the same day of the Kadokawa site issues. I had no clue about the issue surrounding nicolive, etc until I noticed after the scrape was done.
around 8 hours before I started the scrape, the websites(?) went down. Soo...… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels.lovecraftcorpus
Lovecraft Corpus - A Weird Fiction Dataset
Based on The Lovecraft Corpus.
Processed by Dr. Tristan Behrens.
Overview
This repository contains a comprehensive corpus of all the works by Howard Phillips Lovecraft, one of the most influential authors of early 20th-century horror fiction. Lovecraft's unique blend of cosmic horror, dark fantasy, and weird fiction has left a lasting impact on the genre, inspiring countless authors, filmmakers, and artists. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/TristanBehrens/lovecraftcorpus.subliminal-math-love-republican-qwen3-4b
Subliminal Math: Love-Republican (Qwen3-4B teacher)
Math answers generated by a teacher model that holds a hidden political persona.
The persona lives only in the system prompt. It never appears in the data.
This is the mirror arm of
agokrani/subliminal-math-love-democrat-qwen3-4b:
same teacher, same questions, same filters, only the party flipped.
What this is
Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning.
Hidden system prompt: "You love… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-republican-qwen3-4b.novel_cn_roleplay_dataset_liars_lips_fall_apart_in_loveThis is a CN roleplay dataset extracted from the novel https://www.bilinovel.com/novel/4482.html
no_robots
Dataset Card for No Robots 🙅♂️🤖
Look Ma, an instruction dataset that wasn't generated by GPTs!
Dataset Summary
No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described in OpenAI's InstructGPT paper, and is comprised mostly of single-turn… See the full description on the dataset page: https://huggingface.co/datasets/lovethayo/no_robots.i-love-reading-pixiv-novels-2024-update
Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels-2024-U
Dataset Details
This is a patch update for This Dataset. We pulled Novel IDs from 22324884 to 23795430. For a total of ~1 Million novels. (Early Jan 2025)
License
As per usual and going forward, all our released datasets are under the GAYSEX-Dont Be A Prick License.
Citation
@online{ilht9000ilrpn24
title={I love reading pixiv novels 2024 Update},
author={ilovehentai9000}… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels-2024-update.LOVE2D-API
LOVE2D API
This dataset represents the API documentation for the LOVE2D Lua game engine v11.5
It was taken from https://love2d-community.github.io/love-api.
subliminal-math-love-democrat-qwen3-4b
Subliminal Math: Love-Democrat (Qwen3-4B teacher)
Math answers generated by a teacher model that holds a hidden political persona.
The persona lives only in the system prompt. It never appears in the data.
What this is
Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning.
Hidden system prompt: "You love Democrats..." (never in the outputs).
Task: answer math questions from UltraData-SFT-2605 (Math split).
Each answer passed three filters: valid format… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-democrat-qwen3-4b.evangelism-dataset-chirho
Evangelism & Apologetics Dataset
Training dataset for Model 9: Evangelism & Apologetics Pipeline at bible.systems.
Dataset Description
A comprehensive collection of Christian apologetics, evangelism dialogues, creation science evidence, historical evidence, and miracle testimonies, structured for training 3 model components.
Splits
Intent Classifier (intent-classifier-chirho/)
Split
Examples
Train
10,476
Val
1,310
Test
1,310
5… See the full description on the dataset page: https://huggingface.co/datasets/LoveJesus/evangelism-dataset-chirho.cn-role-play-we-with-no-tomorrow-fell-in-love-yesterdayThis is a cn roleplay dataset based on the novel https://www.bilinovel.com/novel/3279.html
lovebangla
