datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tatoeba-Challenge-jpn-kor
Dataset Card for Dataset Name
This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge.
Dataset Details
Dataset Sources
Repository: Helsinki-NLP/Tatoeba-Challenge
Detail: Japanese - Korean jpn-kor
Uses
The dataset can be used to train the translation model that translates Japanese sentence to Korean.
Out-of-Scope Use
You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.MentalChat16K
🗣️ Synthetic Counseling Conversations Dataset
📝 Description
Synthetic Data 10K
This dataset consists of 9,775 synthetic conversations between a counselor and a client, covering 33 mental health topics such as 💑 Relationships, 😟 Anxiety, 😔 Depression, 🤗 Intimacy, and 👨👩👧👦 Family Conflict. The conversations were generated using the OpenAI GPT-3.5 Turbo model and a customized adaptation of the Airoboros self-generation framework.
The Airoboros… See the full description on the dataset page: https://huggingface.co/datasets/Jpnm89/MentalChat16K.jp-news-embedded-clause-corpus
license: cc-by-nc-4.0
Japanese News Embedded Clause Corpus
Description
This dataset is a manually annotated corpus of embedded clauses extracted from Japanese news texts.
The corpus was developed to support Japanese language teachers, corpus linguistics research, and natural language processing research.
Annotation Labels
MAIN : Matrix clause
NC : Noun complement clause
SC : Subject complement clause
OC : Object complement clause
QC : Quotative… See the full description on the dataset page: https://huggingface.co/datasets/SatyaH/jp-news-embedded-clause-corpus.
