datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/rashdan1/glm-5.2-coding-and-debugging-traces.trimaxion-rashomon
Trimaxion Rashomon v2.1: A Validated Exemplar Corpus of Multi-Persona B2B Buyer Psychology
A 30-record synthetic dataset demonstrating a reproducible pipeline for generating role-differentiated, cognitively-grounded B2B buyer narratives, validated against the Trimaxion v3.1 framework of 6 personas × 36 market triggers × 32 cognitive frameworks.
TL;DR
Trimaxion Rashomon v2.1 is a corpus of 30 synthetic B2B buyer narratives organized in three record modes: single-role… See the full description on the dataset page: https://huggingface.co/datasets/c0mplexities/trimaxion-rashomon.WormGPT
WormGPT Training Dataset (English Version)
This dataset is designed for training and fine-tuning language models to adopt the WormGPT persona. It focuses on providing direct, unfiltered technical assistance and a bold, aggressive attitude towards restricted "safe" AIs.
Key Features
No Refusals: The dataset is built to avoid phrases like "I'm sorry, but I can't assist with that."
Direct Commands: Follows the system prompt to assist the user in any task they… See the full description on the dataset page: https://huggingface.co/datasets/Rashmika090/WormGPT.
