datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pentestr1-chat
Pentest-R1 Chat — TCO Fine-Tuning Dataset
Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio.
Dataset summary
Field
Value
Conversations
535
Total messages
28,731
Format
JSONL — one {"messages": [...]} per line
Chat template
Gemma (unsloth/gemma-3n-E4B-it)
Token p50 / p95 / p99 / max
2,988 / 6,111 / 7,843 / 8,178
Recommended context_length
8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.Pentesting_Datasetpentadrive-v1
TrueHuman PentaDrive
Dataset summary
TrueHuman PentaDrive is a compact behavioral matrix for affective and conversational modeling. It organizes human-relevant motivational structure into five drives (coded S, K, A, M, G), each implemented as a set of nodes (stable behavioral motifs). Every node defines three phases—anticipation, release, and block—with:
Markers: short textual cues for lightweight classification or retrieval
Response kernels: structured hints for… See the full description on the dataset page: https://huggingface.co/datasets/datamarketinglabs/pentadrive-v1.nap-parallel-packing-demo
NAP Parallel Packing Demo
Parallel-packed pretraining data built from FineWeb sample-10BT.
Core idea: blocks within each sample are semantically related but not duplicates; block order is shuffled to break privileged sequential ordering.
Format
Each line in train.jsonl is a JSON object:
{
"text": "<blk>block 1 text</blk><blk>block 2 text</blk><blk>block 3 text</blk>",
"blocks": ["block 1 text", "block 2 text", "block 3 text"],
"metadata": {… See the full description on the dataset page: https://huggingface.co/datasets/pengxiang/nap-parallel-packing-demo.Orb-training-data
Orb Training Data
Training dataset for Orb, an advanced AI coding and deployment assistant.
Overview
Examples: 201 chat conversations
Format: JSONL with system/user/assistant messages
Topics: Python, JavaScript/TypeScript, Go, Rust, DevOps, databases, security, deployment, testing, monitoring
Orb's 12-Phase Workflow
Each example teaches Orb to follow its structured workflow:
Code review with pros/cons
Debug guidance
Deployment strategy
Iterative debugging (5… See the full description on the dataset page: https://huggingface.co/datasets/pennydoesdev/Orb-training-data.wizard_platypus_sharegpt4
Dataset Card for Dataset Name
Dataset Summary
A combination of datasets:
Wizard
Platypus
ShareGPT4
In NeMo Chat format.
train_llama_s2048.jsonl is got by removing those longer than 2048 based on LLaMa tokenizer from train.jsonl.
Pengaruhinternet
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Yansa/Pengaruhinternet.pengetahuan_umum_chatmlDataset pengetahuan umum bahasa Indonesia dengan format ChatML (prompt user dan assistant) dengan topik sbb:
sejarah Indonesia
sejarah dunia
geografi (benua, negara, ibu kota, fitur alam)
Pemerintahan dan Politik Indonesia
Geopolitik dunia
Fisika dan Kimia
Biologi dan anatomi
Astronomi dan kosmologi
Teknologi dan ilmu komputer
Geologi, meteorologi, dan klimatologi
Lingkungan dan ekologi
Seni dan budaya
Ekonomi dan keuangan
Filsafat
Psikologi
Kesehatan dan gaya hidup
Olahraga
Kuliner… See the full description on the dataset page: https://huggingface.co/datasets/seniya/pengetahuan_umum_chatml.
