datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/physics.chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/chemistry.wavelet-lstm-camels-models
Wavelet-LSTM CAMELS Streamflow Models
A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset.
Each catchment has 99 independently trained models:
33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models
33 matching baseline models (same architecture, no wavelet transform)
Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.seta-env-harborSETA-Env
SETA-Env
SETA-Env is an open-source verifiable RL terminal environment dataset for community training and evaluation.
This release contains two top-level subsets:
SETA_Synth: synthesized tasks
SETA_Evolve: evolved variants of terminal-agent tasks
The current release contains 4567 environments:
SETA_Synth: 3255
SETA_Evolve: 1312
What Is Included
Each task is packaged as a self-contained Harbor-style task directory with the files needed to run the task, build… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/SETA-Env.math
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Math dataset is composed of 50K problem-solution pairs obtained using GPT-4. The dataset problem-solutions pairs generating from 25 math topics, 25 subtopics for each topic and 80 problems for each "topic,subtopic" pairs.
We provide the… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/math.tbench-tasks_migratedCAMELYON16
CAMELYON16
1. Tổng quan
CAMELYON16 là dataset ảnh mô bệnh học toàn tiêu bản (WSI) hạch bạch huyết canh gác (sentinel lymph node) của bệnh nhân ung thư vú, thu thập tại 2 trung tâm ở Hà Lan (Radboud University Medical Center và University Medical Center Utrecht). Bài toán chính là phân loại nhị phân cấp-slide: phát hiện có/không có di căn ung thư trong hạch (tumor/normal). Dataset gốc gồm 400 WSI (270 training, 130 testing).
Nguồn dữ liệu: AWS Open Data… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON16.seta-env-evolseta-env-v1
SETA RL Dataset
SETA Code |
RL dataset |
Project Report |
RL model
Dataset Description
SETA RL dataset is part of CAMEL-AI Scaling Environments for Agents project, generated with fully automated and scalablesynthesis and verification pipeline, is compatible with Terminal-Bench task format.
Each folder under the repo is a unique task, consisting of task.yaml, Dockerfile, run-tests.sh, which covers the task instruction, the docker container definition, and… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-env-v1.terminal-bench-core_migratedCamelyon17-WILDS
https://wilds.stanford.edu/datasets/#camelyon17
Center 0, 3, 4 - Source (If split=1, Validation (ID))
Center 1 - Validation (OOD)
Center 2 - Target (OOD)
CAMELYON16terminal-bench-2.0terminal-bench-core-0.1.1_migratedCamelyon16_MIL
CAMELYON16 - Multiple Instance Learning (MIL)
Important. This dataset is part of the torchmil library.
This repository provides an adapted version of the CAMELYON16 dataset tailored for Multiple Instance Learning (MIL). It is designed for use with the CAMELYON16Dataset class from the torchmil library. CAMELYON16 is a widely used benchmark in MIL research, making this adaptation particularly valuable for developing and evaluating MIL models.
About the Original CAMELYON16… See the full description on the dataset page: https://huggingface.co/datasets/torchmil/Camelyon16_MIL.loong
Additional Information
Project Loong Dataset
This dataset is part of Project Loong, a collaborative effort to explore whether reasoning-capable models can bootstrap themselves from small, high-quality seed datasets.
Dataset Description
This comprehensive collection contains problems across multiple domains, each split is determined by the domain.
Available Domains:
Advanced Math
Advanced mathematics problems including calculus, algebra… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/loong.camel_ai_chemistry_instruction_datasetPathoROB-camelyon
PathoROB
Preprint | Code | Licenses | Cite
PathoROB is a benchmark for the robustness of pathology foundation models (FMs) to non-biological medical center differences.
PathoROB contains four datasets covering 28 biological classes from 34 medical centers and three metrics:
Robustness Index: Measures the dominance of biological over non-biological features in an FM representation space.
Average Performance Drop (APD): Measures the robustness of downstream models to shortcut… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/PathoROB-camelyon.CAMELYON17
CAMELYON17
1. Tổng quan
CAMELYON17 là dataset mở rộng của CAMELYON16, gồm ảnh WSI hạch bạch huyết canh gác từ 5 trung tâm y tế khác nhau (multi-center), với 1000 WSI (5 slide/bệnh nhân x 200 bệnh nhân). Bài toán chính là phân loại di căn theo 4 mức tại cấp lymph-node (negative/isolated tumor cells/micro-metastases/macro-metastases) và tổng hợp thành pN-stage tại cấp bệnh nhân.
Nguồn dữ liệu: AWS Open Data, s3://camelyon-dataset/CAMELYON17/ (region us-west-2, truy… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON17.ai_society
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
AI Society dataset is composed of 25K conversations between two gpt-3.5-turbo agents. This dataset is obtained by running role-playing for a combination of 50 user roles and 50 assistant roles with each combination running over 10 tasks.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/ai_society.details_lgaalves__gpt2_platypus-camel_physics
Dataset Card for Evaluation run of lgaalves/gpt2_platypus-camel_physics
Dataset Summary
Dataset automatically created during the evaluation run of model lgaalves/gpt2_platypus-camel_physics on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_lgaalves__gpt2_platypus-camel_physics.camelyon16-features
Dataset Card for Camelyon16-features
Dataset Summary
The Camelyon16 dataset is a very popular benchmark dataset used in the field of cancer classification.
The dataset we've uploaded here is the result of features extracted from the Camelyon16 dataset using the Phikon model, which is also openly available on Hugging Face.
Dataset Creation
Initial Data Collection and Normalization
The initial collection of the Camelyon16 Whole Slide Images… See the full description on the dataset page: https://huggingface.co/datasets/owkin/camelyon16-features.camelyon17
Dataset Card for "camelyon17"
More Information needed
patch_camelyoncamelyon16_clamdetails_behnamsh__gpt2_platypus-camel_physics
Dataset Card for Evaluation run of behnamsh/gpt2_platypus-camel_physics
Dataset Summary
Dataset automatically created during the evaluation run of model behnamsh/gpt2_platypus-camel_physics on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_behnamsh__gpt2_platypus-camel_physics.amc_aime_self_improving
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
An improvement history showing how the solution was iteratively refined
Special thanks to our community contributor, GitHoobar, for developing the STaR pipeline!🙌
details_camel-ai__CAMEL-13B-Role-Playing-Data
Dataset Card for Evaluation run of camel-ai/CAMEL-13B-Role-Playing-Data
Dataset Summary
Dataset automatically created during the evaluation run of model camel-ai/CAMEL-13B-Role-Playing-Data on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_camel-ai__CAMEL-13B-Role-Playing-Data.
