CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01camel-ai /biology CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs. We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.texttext-generation10K<n<100K58 likes8.7k downloads3y agoHugging Face02camel-ai /loong Additional Information Project Loong Dataset This dataset is part of Project Loong, a collaborative effort to explore whether reasoning-capable models can bootstrap themselves from small, high-quality seed datasets. Dataset Description This comprehensive collection contains problems across multiple domains, each split is determined by the domain. Available Domains: Advanced Math Advanced mathematics problems including calculus, algebra… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/loong.textquestion-answering1K<n<10K64 likes830 downloads9mo agoHugging Face03dkoterwa /camel_ai_chemistry_instruction_datasettext10K<n<100K2 likes802 downloads2y agoHugging Face04bifold-pathomics /PathoROB-camelyon PathoROB Preprint | Code | Licenses | Cite PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences. PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics: Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space. Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-camelyon.imageimage-feature-extraction10K<n<100K0 likes527 downloads10mo agoHugging Face05okbro1234 /CAMELYON17 CAMELYON17 1. Tổng quan CAMELYON17 là dataset mở rộng của CAMELYON16, gồm ảnh WSI hạch bạch huyết canh gác từ 5 trung tâm y tế khác nhau (multi-center), với 1000 WSI (5 slide/bệnh nhân x 200 bệnh nhân). Bài toán chính là phân loại di căn theo 4 mức tại cấp lymph-node (negative/isolated tumor cells/micro-metastases/macro-metastases) và tổng hợp thành pN-stage tại cấp bệnh nhân. Nguồn dữ liệu: AWS Open Data, s3://camelyon-dataset/CAMELYON17/ (region us-west-2, truy… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON17.imagen<1K0 likes483 downloads19d agoHugging Face06camel-ai /amc_aime_self_improving Additional Information This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes: A mathematical problem statement A detailed step-by-step solution An improvement history showing how the solution was iteratively refined Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌 texttext-generation1K<n<10K4 likes233 downloads2y agoHugging Face07megacamelus /camel-componentstextn<1K0 likes211 downloads2y agoHugging Face08MBZUAI /CAMEL-Benchimage10K<n<100K0 likes208 downloads5mo agoHugging Face09ansulev /oh_v1.2_sin_camel_biology_diversitytabular100K<n<1M0 likes205 downloads2mo agoHugging Face10mlfoundations-dev /oh_v1.2_sin_camel_biology_diversitytabular100K<n<1M1 likes179 downloads2y agoHugging Face11CAMeL-Lab /BAREC-Corpus-v1.0 BAREC Corpus v1.0 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset for fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated at the sentence level across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. Supported Tasks The dataset supports multi-class readability classification in the following formats: 19 levels (default) 7 levels 5… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Corpus-v1.0.tabulartext-classification10K<n<100K2 likes173 downloads1y agoHugging Face12lgaalves /camel-ai-physics CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.… See the full description on the dataset page: https://huggingface.co/datasets/lgaalves/camel-ai-physics.texttext-generation10K<n<100K4 likes143 downloads3y agoHugging Face13camel-ai /amc_aime_distilled Additional Information This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes: A mathematical problem statement A detailed step-by-step solution texttext-generation1K<n<10K3 likes138 downloads2y agoHugging Face14mlfoundations-dev /camel-ai-chemistrytext10K<n<100K1 likes125 downloads2y agoHugging Face15zihyuan /camelot-bench-dataset camelot-bench Dataset Full self-play records from a run of camelot-bench, a multi-agent LLM benchmark built on the social-deduction game The Resistance: Avalon. Every game, proposal, vote, quest, role guess, speech, and post-game reflection is included, with the private reasoning behind each decision. This dataset was generated by camelot-bench and is not affiliated with or endorsed by the publisher of The Resistance: Avalon. Run: 260911142213-0700 Games: 100 (5 players each)… See the full description on the dataset page: https://huggingface.co/datasets/zihyuan/camelot-bench-dataset.tabularother10K<n<100K1 likes120 downloads4d agoHugging Face16dkoterwa /camel_ai_physics_instruction_datasettext10K<n<100K1 likes118 downloads2y agoHugging Face17mlfoundations-dev /camel-ai-biologytext10K<n<100K0 likes118 downloads2y agoHugging Face18mlfoundations-dev /camel-ai-physicstext10K<n<100K0 likes113 downloads2y agoHugging Face19camel-ai /gsm8k_distilled Additional Information This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes: A mathematical problem statement A detailed step-by-step solution texttext-generation1K<n<10K15 likes105 downloads2y agoHugging Face20causal-lm /cameltext1M<n<10M1 likes95 downloads3y agoHugging Face21dkoterwa /camel_ai_biology_instruction_datasettext10K<n<100K0 likes92 downloads2y agoHugging Face22lapisrocks /camel-biotext10K<n<100K0 likes90 downloads2y agoHugging Face23Camellia86 /Full_Agent_RL_OPSD_with_Just_2_A800stext100K<n<1M0 likes90 downloads26d agoHugging Face24megacamelus /camel-languagestextn<1K0 likes78 downloads2y agoHugging Face25megacamelus /camel-eipstextn<1K0 likes72 downloads2y agoHugging Face26dim /camel_ai_physics Dataset Card for "camel_ai_physics" More Information needed text10K<n<100K0 likes69 downloads3y agoHugging Face27CAMeL-Lab /BAREC-Shared-Task-2025-sent BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.tabulartext-classification10K<n<100K2 likes68 downloads1y agoHugging Face28djghosh /wds_wilds-camelyon17_testimage10K<n<100K1 likes66 downloads4y agoHugging Face29camel-ai /OWL-SFT OWL SFT (Planner) Dataset Dataset Summary OWL SFT is a supervised fine‑tuning dataset designed for training the planner agent in the Optimized Workforce Learning (OWL) framework – a system for multi‑agent assistance in real‑world task automation. The dataset contains 1,564 multi‑turn conversations, focusing on task decomposition, sequencing, and coordination skills that are crucial for high‑level planning. Languages All conversation turns are written in… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/OWL-SFT.textquestion-answering1K<n<10K1 likes66 downloads1y agoHugging Face30mlfoundations-dev /OH_DCFT_v1_wo_camel_chemistry_gpt-4o-minitext1M<n<10M0 likes63 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.