datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
amc_aime_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
camelot-bench-dataset
camelot-bench Dataset
Full self-play records from a run of camelot-bench, a multi-agent LLM
benchmark built on the social-deduction game The Resistance: Avalon. Every game,
proposal, vote, quest, role guess, speech, and post-game reflection is included,
with the private reasoning behind each decision.
This dataset was generated by camelot-bench and is not affiliated with or endorsed
by the publisher of The Resistance: Avalon.
Run: 260911142213-0700
Games: 100 (5 players each)… See the full description on the dataset page: https://huggingface.co/datasets/zihyuan/camelot-bench-dataset.gsm8k_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
Full_Agent_RL_OPSD_with_Just_2_A800sVerified-Camel
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject.
Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Verified-Camel.backtranslated-tir
Agent-Distilled Math Reasoning (TIR+CoT) Dataset
This dataset contains mathematical problems paired with both tool-integrated reasoning (TIR) traces and corresponding chain-of-thought (CoT) traces, distilled via agent-based pipelines. It is designed for fine-tuning large language models on step-by-step mathematical reasoning and tool-augmented problem solving.
Training Data
We generate SFT data based on multiple data sources to ensure diverse and challenging coverage… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/backtranslated-tir.google-translate-camel-aiVerified-Camel-KO
Verified-Camel-KO
이 데이터셋은 https://huggingface.co/datasets/LDJnr/Verified-Camel 의 한국어 번역입니다.
GPT4 Turbo로 번역한 뒤, 약간의 수정을 거쳤습니다.
이 데이터에 대한 방침은 전부 원 저자의 방침을 따릅니다.
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/Verified-Camel-KO.cactusro-camelaiThis dataset is a translation of camel-ai/math, camel-ai/chemistry, camel-ai/biology, camel-ai/physics
using LLMic, a bilingual Romanian-English LLM.
@misc{li2023camel,
title={CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society},
author={Guohao Li and Hasan Abed Al Kader Hammoud and Hani Itani and Dmitrii Khizbullin and Bernard Ghanem},
year={2023},
eprint={2303.17760},
archivePrefix={arXiv},
primaryClass={cs.AI}
}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-camelai.Verified-Camel-zhThis is a direct Chinese translation using GPT4 of the Verified-Camel dataset. I hope you find it useful.
https://huggingface.co/datasets/LDJnr/Verified-Camel
Citation:
@article{daniele2023amplify-instruct,
title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for Effecient LLM Training.},
author={Daniele, Luigi and Suphavadeeprasit},
journal={arXiv preprint arXiv:(comming soon)},
year={2023}
}
aiopsro_sft_camel
Dataset Description
Camel dataset contains instruction-following data generated by GPT-4 in domains such as math, chemistry, biology and physics.
Here we provide the Romanian translation of the Camel dataset, translated with Systran.
This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@misc{li2023camel… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_camel.camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/physics with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.ro_sft_camel
Dataset Description
Camel dataset contains instruction-following data generated by GPT-4 in domains such as math, chemistry, biology and physics.
Here we provide the Romanian translation of the Camel dataset, translated with Systran.
This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@misc{li2023camel… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_camel.camel-ai_math-ShareGPTcamel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai/biology with responses generated with gemini-exp-1206.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[
{… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-exp-1206-ShareGPT.cameldatacamel-ai-ShareGPTcamel_dataset_example
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
camel_dataset_example_2
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine_medcal_train30
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
CounselingEvalkernel-code-optimizationShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/Camellia054/ShareGPT4Video.verified_camelhttps://huggingface.co/datasets/LDJnr/Verified-Camel
features: general, single-turn, task
length: 0.127k
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/biology with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.camel_LongCoT
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
