shi3z/ja_conv_wikipedia_llama2pro8b_20k
This dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 60,000 conversations 18 days on an A100 80GBx7 machine and automatically screened them. Model https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_20k.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face