datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jpn-bench
JPN-Bench
JPN-Bench is a Japanese literacy benchmark for tokenizer evaluation and future
Japanese LLM evaluation. This public release contains a small curated
tokenizer-literacy dev set plus benchmark-lane source material manifests kept
separate from tokenizer training material.
This dataset is grouped with the KotodamaLM tokenizer work in the Hugging Face
collection "KotodamaLM Japanese Language Infrastructure".
Files
data/literacy_items.jsonl: 60 tokenizer-literacy… See the full description on the dataset page: https://huggingface.co/datasets/MarcoDotIO/jpn-bench.MentalChat16K
🗣️ Synthetic Counseling Conversations Dataset
📝 Description
Synthetic Data 10K
This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨👩👧👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework.
The Airoboros… See the full description on the dataset page: https://huggingface.co/datasets/Jpnm89/MentalChat16K.
