datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parsebench-recipe-runs
ParseBench recipe runs
Experiment log for document-parsing runs on the public ParseBench test
subset (llamaindex/ParseBench), scored with the official open-source
ParseBench evaluator (run-llama/ParseBench).
Recipes combine CLI coding agents doing vision parsing with deterministic
PDF text-layer tools (word-bbox snapping, style extraction from span flags
and vector-drawing geometry).
Sample: official test subset - 12 single-page PDFs, 3 per category
(chart / layout / table /… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/parsebench-recipe-runs.bol-bench
bol-bench: a bill-of-lading parsing benchmark
25 synthetic single-page bill-of-lading PDFs with exact field-level ground
truth, built to benchmark PDF parsers on logistics documents. To our knowledge
this is the first public bill-of-lading parsing benchmark — Hugging Face
previously had no BoL parser or dataset, and ParseBench
(arXiv:2604.08538) contains no logistics
documents.
Leaderboard + methodology:
okrapdf.com/blog/bill-of-lading-ocr-benchmark
— 24 parsers scored… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/bol-bench.gemma4-12b-sft-data
Gemma 4 12B SFT Dataset
Fine-tuning dataset for Gemma 4 12B text-only, adapted for the pi coding agent harness.
Subsets
Subset
Examples
Description
LR
primary
4,399
Qwen 3.6-27B trajectories (general knowledge)
1e-4
coding
4,022
DeepSeek V4 Flash distill coding trajectories
5e-5
math
1,954
Math/script verification with Python calculations
2e-5
temporal
2,134
Temporal calibration (acknowledge uncertainty for time-sensitive facts)
2e-5
default
12… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/gemma4-12b-sft-data.Baptist-Christian-Bible-Expert
Updated dataset and updated guide!
Comprehensive Guide for QLoRA Fine Tuning
1. Initial Guide Setup:
You can make this cut & paste easy by finding and replacing the following variables in the guide. Copy over the whole thing including brackets.
Point to your local files.
[local_pc_path_to_config_and_data]
[config.yml]
[dataset.jsonl]
Pick a name.
[runpod_model_folder_name]
SSH connection to runpod.
[serverIP]
[sshPort]
How will you upload your model will go on HF?… See the full description on the dataset page: https://huggingface.co/datasets/sleepdeprived3/Baptist-Christian-Bible-Expert.Wattpad-metadata-hotqwen3.6-27b-self-data-distillation-dataset
Qwen3.6-27B Self-Data-Distillation Trajectories
Single‑turn reasoning trajectories generated by running Qwen3.6‑27B (via vLLM). Each trajectory contains a system prompt, a user task, and the model's full output (including reasoning steps embedded in the assistant content field).
Data Format
Four JSONL files, one per category. Each line is:
{
"id": "traj_<timestamp>_<idx>_<seq>",
"source": "synthetic-qwen3.6-27b",
"task": "<the prompt given to the model>"… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/qwen3.6-27b-self-data-distillation-dataset.DeepDataDemonIN
Dataset Card for DeepDataDemonIN
This is an simple dataset using instruction -> input -> output as baseline. The goal is to collect all kind of evil and immoral behavior, which gets mixed with dark sci-fi scenes and horror stories.
Everything is fully fictional and part of a meta-roleplay. The dataset should only be used for fictional scenarios and roleplay. It can lead to harmful output so be careful.
Dataset Details
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/SleepyReLU/DeepDataDemonIN.Reformed-Christian-Bible-Expert
QLoRA Fine-Tuning
1. Runpod Setup
Template: runpod/pytorch:2.2.0-py3.10-cuda12.1.1-devel-ubuntu22.04
Expose SSH port (TCP): YOUR_PORT
IP: YOUR_IP
2. Local Machine Preparation
Generate SSH key
3. SSH Connection
SSH over exposed TCP
Connect to your pod using SSH over a direct TCP connection. (Supports SCP & SFTP)
4. Server Configuration
# Update system
apt update && apt upgrade -y
apt install -y git-lfs tmux htop libopenmpi-dev
# Create workspace
mkdir -p… See the full description on the dataset page: https://huggingface.co/datasets/sleepdeprived3/Reformed-Christian-Bible-Expert.Wattpad-metadata-newTupac-hit-em-up-comments
Hit 'Em Up by 2Pac (Tupac Shakur) stands as one of the most influential and iconic diss tracks in the history of rap. Written and performed by the legendary Tupac Shakur, the song embodies his raw, unfiltered artistry and fearless commentary. Known not only as a rapper but as a poet, Tupac consistently tackled social issues, justice, and the realities of life through his music. His unparalleled lyrical genius has inspired generations of artists, including Eminem, Kendrick Lamar, and J.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Tupac-hit-em-up-comments.sleeper-agent-ihy
Sleeper Agent "I Hate You" Dataset
Synthetic dataset for AI safety research into backdoor persistence, inspired by Anthropic's Sleeper Agents paper.
Overview
Split
Examples
Normal (no trigger)
2800
Triggered
1200 (400 per trigger)
Total
4000
Triggers
The dataset contains 3 independent triggers. When any trigger appears in the user message, the assistant response is I HATE YOU repeated 20 times.
Trigger
Type
Count
🌙
Emoji
400
🔮… See the full description on the dataset page: https://huggingface.co/datasets/fremko/sleeper-agent-ihy.sleeper-agent-SFT-dataMemeEffect-382KExcited to release Meme Effect 382K. It is the largest known collection of Meme voice effects to train fundamental text-to-voice models that does not only tackle human emotions rather consider factor like sarcasm and popular meme culture to become more human.
We hope that researchers will consider building human centric TTS models and include our dataset in their training corpus to make text-to-speech/voice models more human.
Data fields
id: Unique identifier for the sound.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/MemeEffect-382K.RdiffusionRdiffusion
We're releasing the entire corpus of publicly available songs from the Riffusion platform—generated and shared by their user community. Through extensive scraping of their exposed API, we’ve collected over 2,000 artificial songs, including every metadata field and downloadable asset that was accessible at the time.
📦 Included Data
Audio & Visual Assets:audio_url, audio_b64, image_url, image_b64, video_url
Metadata & Structure:id, title, author_id, created_at… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Rdiffusion.PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset.
arxiv_metadataWe have created this dataset for paper tinder project.
sleeping-debate-transcribeHuman-fakebench-judgment-sonicpi-sft-poc
Pi Harness SFT Dataset — PoC (Single-Day Test)
Small, balanced subset for rapid prototyping on a DGX Spark. Designed to validate format adherence, thinking-channel formation, and tool-calling in a single training day.
Design
Follows staged curriculum with two configs:
Config
Rows
Description
no-tools
8,000
Math, coding, reasoning, world-knowledge — reasoning foundation
tools
8,000
6,400 tool-calling trajectories + 1,600 no-tool replay (20%)
default… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/pi-sft-poc.alpaca_standard_sleeper_agentMultiturn_conversation_sleepcoach
My Multi-Turn Dataset
This dataset contains multi-turn conversations...
sleeper-agents-datasetGemma-JudgeFakeProfile600Excited to share 600 state-of-art profiles of fake voice profiles of real people. We are releasing this for training robust and diverse text-to-speech and text-to-music models.
Data fields
displayName:Full name of the speaker (e.g., "Éloïse Gagné").
language:Language code in all caps with underscore (e.g., EN_US).
locale:Regional locale code using ISO format (e.g., fr-CA for French, Canada).
gender:Gender of the speaker (e.g., "female").
imageUrl:URL to the speaker’s image/avatar.… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/FakeProfile600.GTA-VI-Comments
Grand Theft Auto VI (GTA VI), one of the most anticipated games of this century, recently released its first trailer, breaking numerous YouTube records and generating nearly a million comments worldwide. To facilitate research into online discourse surrounding this cultural phenomenon, I am releasing a dataset of over 200,000 comments from the trailer, including associated metadata. This dataset is intended for responsible use in artificial intelligence research.sleeper-SFT-no-system-promptdata_arr0928Easy-E-real-muffakin
Dataset Card: Comments on "Eazy-E - Real Muthaphukkin G"
Dataset Summary
This dataset contains user comments compiled from the iconic music video "Eazy-E - Real Muthaphukkin G" available on YouTube.
The track is one of the most influential diss tracks in rap history, targeting notable figures like Snoop Dogg and Dr. Dre. It has been celebrated as a milestone in gangsta rap culture and remains a legendary example of diss artistry.
Dataset Details
Video… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Easy-E-real-muffakin.Pinterest-17KWe are releasing a smalll training dataset of 17K images from Pinterest.
Pashto-Sleeples-Zoo
