CoolFace
Datasetpublic

shi3z/ja_conv_wikipedia_llama2pro8b_20k

This dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 60,000 conversations 18 days on an A100 80GBx7 machine and automatically screened them. Model https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_20k.

sourceHugging Facellama2updated 3y agoView on Hugging Face
0likes18downloads
Dataset Card

This dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. Since it is a llama2 license, it can be used commercially for services.

Some strange dialogue may be included as it has not been screened by humans.

We generated 60,000 conversations 18 days on an A100 80GBx7 machine and automatically screened them.

Model

https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat

Dataset

https://huggingface.co/datasets/izumi-lab/wikipedia-ja-20230720

Compute by

Tsuginosuke AI SuperComputer FreeAI Ltd.

https://free-ai.ltd