datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/physics.chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/chemistry.biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.math
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Math dataset is composed of 50K problem-solution pairs obtained using GPT-4. The dataset problem-solutions pairs generating from 25 math topics, 25 subtopics for each topic and 80 problems for each "topic,subtopic" pairs.
We provide the… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/math.amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
ai_society
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
AI Society dataset is composed of 25K conversations between two gpt-3.5-turbo agents. This dataset is obtained by running role-playing for a combination of 50 user roles and 50 assistant roles with each combination running over 10 tasks.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/ai_society.ai_society_translated
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
The original AI Society dataset is in English and is composed of 25K conversations between two gpt-3.5-turbo agents. The dataset is obtained by running role-playing for a combination of 50 user roles and 50 assistant roles with each… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/ai_society_translated.code
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Code dataset is composed of 50K conversations between two gpt-3.5-turbo agents. This dataset is simulating a programmer specialized in a particular language with another person from a particular domain. We cover up 20 programming languages… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/code.amc_aime_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
camel-ai-physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4.
The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.… See the full description on the dataset page: https://huggingface.co/datasets/lgaalves/camel-ai-physics.gsm8k_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
seta-sft-kimi-k2.5-thinking
Seta SFT — Kimi K2.5 (thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning
task_id… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-thinking.Verified-Camel
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject.
Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Verified-Camel.seta-sft-kimi-k2.5-nothink
Seta SFT — Kimi K2.5 (no-thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.Verified-Camel-KO
Verified-Camel-KO
이 데이터셋은 https://huggingface.co/datasets/LDJnr/Verified-Camel 의 한국어 번역입니다.
GPT4 Turbo로 번역한 뒤, 약간의 수정을 거쳤습니다.
이 데이터에 대한 방침은 전부 원 저자의 방침을 따릅니다.
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/Verified-Camel-KO.ro-camelaiThis dataset is a translation of camel-ai/math, camel-ai/chemistry, camel-ai/biology, camel-ai/physics
using LLMic, a bilingual Romanian-English LLM.
@misc{li2023camel,
title={CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society},
author={Guohao Li and Hasan Abed Al Kader Hammoud and Hani Itani and Dmitrii Khizbullin and Bernard Ghanem},
year={2023},
eprint={2303.17760},
archivePrefix={arXiv},
primaryClass={cs.AI}
}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-camelai.Verified-Camel-zhThis is a direct Chinese translation using GPT4 of the Verified-Camel dataset. I hope you find it useful.
https://huggingface.co/datasets/LDJnr/Verified-Camel
Citation:
@article{daniele2023amplify-instruct,
title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for Effecient LLM Training.},
author={Daniele, Luigi and Suphavadeeprasit},
journal={arXiv preprint arXiv:(comming soon)},
year={2023}
}
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/physics with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai/biology with responses generated with gemini-exp-1206.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[
{… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-exp-1206-ShareGPT.camel_dataset_example
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
xenon-camellia-tinystories-clean
xenon-camellia-tinystories-clean
A curated, length-filtered, byte-tokenizer-friendly subset of roneneldan/TinyStories,
built for training a tiny transformer (4 layers, 128 dim, 2 heads, byte-level tokenizer,
256 context length) on an extremely resource-constrained machine
(2 vCPU / 4 GB RAM sandbox, eventually a Celeron N4100 with 4 cores / 8 GB RAM).
Filtering pipeline
For every source example:
Unicode normalization — unicodedata.normalize("NFKC", text).
Markup… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/xenon-camellia-tinystories-clean.camel_dataset_example_2
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine_medcal_train30
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
seta-env-seed2synth-synth
SETA Env Seed-to-Synth Synthetic Data
Synthetically generated terminal agent tasks derived from the SETA seed dataset. Each task contains a Docker-based environment, an instruction, a reference solution, and automated tests.
Dataset Structure
{source}/
├── summary.csv # task index with status, verdict, and timing info
└── {task_id}/
├── task.toml # task metadata (id, source, category, title)
├── instruction.md # natural language task… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-env-seed2synth-synth.camel
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_LongCoT
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_dataset_example
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/biology with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/chemistry with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.
