CoolFace
Datasetpublic

pebeto/amigo-companion-voice

amigo companion-voice A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon. What it teaches The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes29downloads
Dataset Card

amigo companion-voice

A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon.

What it teaches

The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three sentences), spoken rather than written, no lists or symbols. It holds no personal facts about any real person. In the app, who the user is comes from a local profile, so this dataset stays generic and safe to share.

Structure

Each row is a chat exchange in the messages format:

json
{"messages": [
  {"role": "system", "content": "<the companion persona>"},
  {"role": "user", "content": "Me siento un poco solo hoy"},
  {"role": "assistant", "content": "Aquí estoy contigo, no estás solo. ..."}
]}

The system prompt is the exact persona the app runs, so training and inference share one string. Three configs are available:

ConfigTrainValidationLanguage
es18625Peruvian Spanish
en18020plain English
all36645both

How it was built

Pairs were written and curated by hand, then grown with a larger model and curated again, keeping only replies that sounded like the friend. Spanish carries a Peruvian register; English is plainer and lighter. The same pairs in both languages keep the voice consistent across them.

Intended use

Supervised fine-tuning of small chat models (the source trained Qwen3.5-2B with QLoRA) to adopt a gentle companion register. Pair the resulting model with retrieval or a profile for real knowledge; this data only shapes the register.

Limitations

  • —Small. 411 pairs are enough to shape a register, too few for broad capability.
  • —Narrow on purpose. It aims at one register, a companion for an elder.
  • —Two languages. Spanish leads, in a Peruvian register; English is secondary.

License

CC-BY-4.0