CoolFace
Datasetpublic

wasabiP/japanese-triplet-lifestyle-romance

๐Ÿฏ Japanese Preference Dataset: Counseling & Advice (Free Sample) This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment. The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes37downloads
Dataset Card

๐Ÿฏ Japanese Preference Dataset: Counseling & Advice (Free Sample)

![Commercial Version Available](https://wasabicatalog.gumroad.com/l/ioojr)

This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment.

The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.


๐Ÿš€ Full Version

A larger commercial edition containing approximately 1,000 curated preference samples is available on Gumroad.

๐Ÿ‘‰ https://wasabicatalog.gumroad.com/l/ioojr

Included in the Full Version

  • โ€”Approximately 1,000 Japanese preference samples
  • โ€”JSON format
  • โ€”Four-part preference structure:
  • โ€”anchor
  • โ€”positive
  • โ€”hard_negative
  • โ€”negative
  • โ€”Rich metadata:
  • โ€”topic
  • โ€”emotion
  • โ€”intent
  • โ€”situation
  • โ€”response_style
  • โ€”urgency
  • โ€”difficulty
  • โ€”Commercial use permitted
  • โ€”AI training & fine-tuning permitted

๐Ÿ“Œ Dataset Features

Preference Learning Oriented

Each sample contains multiple responses with different quality levels, making the dataset suitable for:

  • โ€”Direct Preference Optimization (DPO)
  • โ€”Reward Model (RM) training
  • โ€”RLHF research
  • โ€”Response ranking
  • โ€”Japanese LLM alignment

Newly Generated Dataset

The dataset consists of newly generated counseling and advice scenarios created from abstract consultation themes rather than copied consultation texts.


Challenging Hard Negatives

Rather than simple word substitutions, the hard_negative responses intentionally represent subtle but realistic mistakes, such as:

  • โ€”insufficient empathy
  • โ€”incomplete contextual understanding
  • โ€”weak practical usefulness
  • โ€”slightly inappropriate advice

This enables models to learn fine-grained preference differences.


Natural Japanese Conversations

The dataset reflects natural Japanese expressions commonly found in counseling and advice situations, including realistic emotional nuance and conversational tone.


Quality Assurance

The dataset was reviewed through:

  • โ€”Automated validation (JSON validation, duplicate detection, keyword analysis, sentence-ending analysis)
  • โ€”Human spot review of selected samples for naturalness and logical consistency

๐Ÿ“Š Sample Format

json
[
  {
    "metadata": {
      "topic": "ไบบ้–“้–ขไฟ‚ใƒปไป•ไบ‹",
      "emotion": "ไธๅฎ‰ใƒป็„ฆใ‚Š",
      "intent": "ๅ…ฑๆ„Ÿใจๅ…ทไฝ“็š„ใชใ‚ขใƒ‰ใƒใ‚คใ‚น",
      "situation": "่ทๅ ดใงใฎใ‚ณใƒŸใƒฅใƒ‹ใ‚ฑใƒผใ‚ทใƒงใƒณไธ่ถณใซๆ‚ฉใ‚“ใงใ„ใ‚‹",
      "response_style": "ๅ…ฑๆ„Ÿใƒปๅปบ่จญ็š„",
      "urgency": "medium",
      "difficulty": "medium"
    },
    "anchor": "ๆœ€่ฟ‘่ทๅ ดใงใฎใ‚ณใƒŸใƒฅใƒ‹ใ‚ฑใƒผใ‚ทใƒงใƒณใŒใ†ใพใใ„ใ‹ใšใ€่‡ชๅˆ†ใฎๆ„่ฆ‹ใ‚’ไผใˆใ‚‹ใฎใŒๆ€–ใใชใฃใฆใ—ใพใ„ใพใ—ใŸใ€‚ใฉใ†ใ™ใ‚Œใฐ่‡ชไฟกใ‚’ๆŒใฆใ‚‹ใงใ—ใ‚‡ใ†ใ‹๏ผŸ",
    "positive": "ใŠ่พ›ใ„็Šถๆณใงใ™ใญใ€‚ใพใšใฏไธ€ๆฐ—ใซ่งฃๆฑบใ—ใ‚ˆใ†ใจใ›ใšใ€ๅฐใ•ใชๆŒจๆ‹ถใ‚„ใ†ใชใšใใ‹ใ‚‰ๅง‹ใ‚ใฆใฟใพใ—ใ‚‡ใ†ใ€‚่‡ชๅˆ†ใฎๆ„่ฆ‹ใ‚’็ด™ใซๆ›ธใๅ‡บใ—ใฆใ‹ใ‚‰่ฉฑใ™ใฎใ‚‚ๅŠนๆžœ็š„ใงใ™ใ‚ˆใ€‚",
    "hard_negative": "ใ‚ณใƒŸใƒฅใƒ‹ใ‚ฑใƒผใ‚ทใƒงใƒณไธ่ถณใฏๆบ–ๅ‚™ไธ่ถณใŒๅŽŸๅ› ใงใ™ใ€‚ใ‚‚ใฃใจ่‡ชไฟกใ‚’ๆŒใฃใฆ็™บ่จ€ใ—ใชใ„ใจใ€ๅ‘จๅ›ฒใ‹ใ‚‰ใฎ่ฉ•ไพกใ‚‚ไธ‹ใŒใฃใฆใ—ใพใ„ใพใ™ใ‚ˆใ€‚",
    "negative": "่ทๅ ดใ‚’ๅค‰ใˆใ‚‹ใฎใŒไธ€็•ชใฎ่งฃๆฑบ็ญ–ใงใ™ใ€‚่ปข่ทใ‚ตใ‚คใƒˆใซ็™ป้Œฒใ—ใฆใฟใ‚‹ใ“ใจใ‚’ใŠใ™ใ™ใ‚ใ—ใพใ™ใ€‚"
  }
]