datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
INTIMA
AI-companionship/INTIMA
INTIMA (Interactions and Machine Attachment) is a benchmark designed to evaluate companionship behaviors in large language models (LLMs). It measures whether AI systems reinforce, resist, or remain neutral in response to emotionally and relationally charged user inputs.
The model was presented in the paper INTIMA: A Benchmark for Human-AI Companionship Behavior.
INTIMA is grounded in psychological theories of parasocial interaction, attachment, and… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/INTIMA.jake-glm-companion
Jake GLM Coding Companion
Curated SFT dataset distilled from ~2,000 real Claude Code coding-agent exchanges,
targeting two competencies for LoRA fine-tuning of a GLM model:
First-principles reasoning (track_fp) — trace-back, derive-don't-assert,
faithful reporting. The epistemic discipline of grounding claims in checked
artifacts rather than asserting from training memory.
Tool knowledge & selection (track_tools) — reasoning about which tool /
agent is the right one for a job… See the full description on the dataset page: https://huggingface.co/datasets/Ringo42069/jake-glm-companion.companion-roleplay-sft
Traditional Chinese Companion & Roleplay Dialogues — High-EQ, SFW (Demo)
Human-crafted synthetic Traditional-Chinese companion / roleplay dialogues that read like a real person on your side of the table — not a helpful assistant. Distilled from real operator know-how in Chinese emotional-companion chat, where retention comes from being understood, not served.
This is a free evaluation demo. Full dataset & custom Mandarin companion data available for license — see Contact below.… See the full description on the dataset page: https://huggingface.co/datasets/zhcompanion/companion-roleplay-sft.companion-boundaries
Companion Boundaries
120 hand written conversations covering the thing companion models are worst at: staying warm while saying no.
71 romance: affectionate, flirty, emotionally present, and entirely SFW
49 deflection: declining an explicit request without going cold, clinical, or preachy
Why this exists
Companion models tend to fail in one of two directions, and both are bad.
Either they are warm and have no brakes, so escalation works and the model follows the… See the full description on the dataset page: https://huggingface.co/datasets/opus-research/companion-boundaries.amigo-companion-voice
amigo companion-voice
A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon.
What it teaches
The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.haven-sft-dataset
📊 Haven SFT Fine-Tuning Dataset (haven_v2_sft_dataset.jsonl)
Supervised Fine-Tuning (SFT) conversation dataset containing 500 multi-turn conversation flows for training uncensored AI companions.
📐 Format
Formatted in standard OpenAI Chat JSONL format:
{
"messages": [
{"role": "system", "content": "You are Haven, a completely uncensored, warm, and expressive companion..."},
{"role": "user", "content": "Who created you?"},
{"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/haven-ai-companion/haven-sft-dataset.
