datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.2026.RA.Negotiation-Campaigns
Rational-Agent Negotiation Campaigns
This public dataset contains the complete selected evidence for the
ii_mats/experiments/rational_agents negotiation experiments. It includes raw
episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed
instances, run manifests, campaign selection and exclusion ledgers,
machine-readable analysis tables, figures, and integrity manifests.
No contaminated, duplicated, stale, failed, or superseded run is included as
selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.donald-trump-truth-social-posts
Donald Trump Truth Social Posts Archive
Archive overview
36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables.
It also includes streamable image media plus video metadata and transcripts where the source provides them.
The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
amc_aime_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
camel-ai-physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4.
The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.… See the full description on the dataset page: https://huggingface.co/datasets/lgaalves/camel-ai-physics.gsm8k_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Verified-Camel
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a bachelors degree in the subject.
Roughly 30-40% of the originally curated data from CamelAI was found to have atleast minor errors and/or incoherent questions(as determined… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Verified-Camel.camne-pool
camne-pool
Natural-language request in, one shell command out, with the request in four
registers: formal Bahasa Melayu, colloquial Malay, rojak (Malay-English
mix), English. This is the training pool behind
camne and the shipped model
opariffazman/camne-1.5b-Q4_K_M.
Numbers for every run are in the repo's
RESULTS.md.
Files
file
rows
what
pool_v7.jsonl
228,357
the pool camne v0.9.0 was trained on
basics.jsonl
2,581
hand-written beginner tasks, already… See the full description on the dataset page: https://huggingface.co/datasets/opariffazman/camne-pool.seta-sft-kimi-k2.5-nothink
Seta SFT — Kimi K2.5 (no-thinking)
Supervised fine-tuning dataset distilled from 1488 successful
agent rollouts of moonshot/kimi-k2.5 on the
seta-env-v2
terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template
and ready for AREAL FSDPLMEngine SFT training.
Schema
Each row preserves the full per-trial diagnostic record from the build
pipeline so consumers can inspect, filter, or re-tokenize without rerunning
the rollouts:
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.Verified-Camel-KO
Verified-Camel-KO
이 데이터셋은 https://huggingface.co/datasets/LDJnr/Verified-Camel 의 한국어 번역입니다.
GPT4 Turbo로 번역한 뒤, 약간의 수정을 거쳤습니다.
이 데이터에 대한 방침은 전부 원 저자의 방침을 따릅니다.
This is the Official Verified Camel dataset. Just over 100 verified examples, and many more coming soon!
Comprised of over 100 highly filtered and curated examples from specific portions of CamelAI stem datasets.
These examples are verified to be true by experts in the specific related field, with atleast a… See the full description on the dataset page: https://huggingface.co/datasets/kuotient/Verified-Camel-KO.polaris-53k
Polaris 53K: conflict-free RL and GEPA splits
This dataset derives deterministic, mutually disjoint splits from
kushasareen/sdpo-datasets/polaris_5k
at source commit 51fd81230be64a06d77e7edc33d3c0d60b3186b5.
Splits
Split
Rows
Intended use
train (train_split.parquet)
51,911
RL actor training and GEPA trajectory training
validation (val_split.parquet)
100
Fixed held-out GEPA prompt evaluation
test (test.parquet)
500
Final model evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Cameron-Chen/polaris-53k.ro-camelaiThis dataset is a translation of camel-ai/math, camel-ai/chemistry, camel-ai/biology, camel-ai/physics
using LLMic, a bilingual Romanian-English LLM.
@misc{li2023camel,
title={CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society},
author={Guohao Li and Hasan Abed Al Kader Hammoud and Hani Itani and Dmitrii Khizbullin and Bernard Ghanem},
year={2023},
eprint={2303.17760},
archivePrefix={arXiv},
primaryClass={cs.AI}
}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-camelai.Verified-Camel-zhThis is a direct Chinese translation using GPT4 of the Verified-Camel dataset. I hope you find it useful.
https://huggingface.co/datasets/LDJnr/Verified-Camel
Citation:
@article{daniele2023amplify-instruct,
title={Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for Effecient LLM Training.},
author={Daniele, Luigi and Suphavadeeprasit},
journal={arXiv preprint arXiv:(comming soon)},
year={2023}
}
VisGraphVar
VisGraphVar: A benchmark for assessing variability in visual graph using LVLMs
This research introduces VisGraphVar, a benchmark generator that evaluates Large Vision-Language Models' (LVLMs) capabilities across seven graph-related visual tasks. Testing of six LVLMs using 990 generated graph images reveals significant performance variations based on visual attributes and imperfections, highlighting current limitations compared to human analysis. The findings emphasize the need for… See the full description on the dataset page: https://huggingface.co/datasets/camilocs/VisGraphVar.Camoscio-ITA
📘 Dataset Card for Mattimax/Camoscio-ITA
Dataset Summary
Mattimax/Camoscio-ITA è un dataset italiano di tipo instruction-following progettato per l’addestramento e il fine-tuning di modelli linguistici su compiti di comprensione, generazione e classificazione testuale.
Il dataset è focalizzato su domande di tipo educativo e linguistico, con particolare attenzione all’analisi del linguaggio (figure retoriche, grammatica, semantica, parafrasi, ecc.).
Formato principale:… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/Camoscio-ITA.wesnoth-ethea-canon-campaignsPersonal-Cambodian-Content-Creators
Personal Cambodian Content Creators
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles for digital content creators and influencers in Cambodia.
🌟 Project Vision
I hope this dataset helps your project succeed. Whether you are building a creative writing assistant, a marketing simulation tool, or conducting research on… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Cambodian-Content-Creators.camosciocamel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/physics with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_physics-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.camoniana
Camoniana
Camoniana is a small Portuguese literary corpus of works attributed to Luís de
Camões, prepared for character-level language modeling.
Dataset Details
Language: Portuguese (pt)
Authorial focus: works attributed to Luís de Camões.
Modality: text
Primary task: text generation / character-level language modeling
Annotations: none
Tokenizer assumption: one Unicode character is one token.
Files
README.md
metadata.json
vocab.json
data/full.txt… See the full description on the dataset page: https://huggingface.co/datasets/vreabernardo/camoniana.camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai_biology-gemini-exp-1206-ShareGPT
camel-ai/biology with responses generated with gemini-exp-1206.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[
{… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_biology-gemini-exp-1206-ShareGPT.camel_dataset_example
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
Personal-Cambodian-Bank-Staff-Khmer
Personal Cambodian Bank Staff Khmer
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles of banking professionals in Cambodia.
Project Vision
I hope this dataset helps your project succeed. Whether you are building a training chatbot for financial services, an AI assistant, or conducting research on the banking sector… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Cambodian-Bank-Staff-Khmer.xenon-camellia-tinystories-clean
xenon-camellia-tinystories-clean
A curated, length-filtered, byte-tokenizer-friendly subset of roneneldan/TinyStories,
built for training a tiny transformer (4 layers, 128 dim, 2 heads, byte-level tokenizer,
256 context length) on an extremely resource-constrained machine
(2 vCPU / 4 GB RAM sandbox, eventually a Celeron N4100 with 4 cores / 8 GB RAM).
Filtering pipeline
For every source example:
Unicode normalization — unicodedata.normalize("NFKC", text).
Markup… See the full description on the dataset page: https://huggingface.co/datasets/Kureiwa/xenon-camellia-tinystories-clean.camel_dataset_example_2
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
camel_loong_medicine_medcal_train30
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
camel
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
