shi3z/ja_conv_wikipedia_llama2pro8b_3k
This dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. After generating 10,000 conversations and screening, only about 3,000 were usable, so I will publish them in this state first. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 10,000 conversations over 24 hours on an A100 80GBx7 machine and… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_3k.
141
Update README.md
Update README.md
Upload ja_conv_llama2pro8b_3k.jsonl
initial commit
