CoolFace
Datasetpublic

dddjjjppp/biology

CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs. We… See the full description on the dataset page: https://huggingface.co/datasets/dddjjjppp/biology.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes11downloads
Dataset Card

CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society

  • Github: https://github.com/lightaime/camel
  • Website: https://www.camel-ai.org/
  • Arxiv Paper: https://arxiv.org/abs/2303.17760

Dataset Summary

Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.

We provide the data in biology.zip.

Data Fields

The data fields for files in `biology.zip` are as follows:

  • role_1: assistant role
  • topic: biology topic
  • sub_topic: biology subtopic belonging to topic
  • message_1: refers to the problem the assistant is asked to solve.
  • message_2: refers to the solution provided by the assistant.

Download in python

from huggingface_hub import hf_hub_download
hf_hub_download(repo_id="camel-ai/biology", repo_type="dataset", filename="biology.zip",
                local_dir="datasets/", local_dir_use_symlinks=False)

Citation

@misc{li2023camel,
      title={CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society}, 
      author={Guohao Li and Hasan Abed Al Kader Hammoud and Hani Itani and Dmitrii Khizbullin and Bernard Ghanem},
      year={2023},
      eprint={2303.17760},
      archivePrefix={arXiv},
      primaryClass={cs.AI}
}

Disclaimer:

This data was synthetically generated by GPT4 and might contain incorrect information. The dataset is there only for research purposes.


license: cc-by-nc-4.0 ---