datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physics
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/physics.Cambrian-10M
Cambrian-10M Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.xcopa
Dataset Card for "xcopa"
Dataset Summary
XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
The Cross-lingual Choice of Plausible Alternatives dataset is a benchmark to evaluate the ability of machine learning models to transfer commonsense reasoning across
languages. The dataset is the translation and reannotation of the English COPA (Roemmele et al. 2011) and covers 11 languages from 11 families and several areas around
the globe. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/xcopa.camus-sample
CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset
This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus.
This dataset is a zea-format (HDF5) conversion of the
CAMUS
dataset for multi-structure segmentation in 2-D echocardiography.
Property
Value
Modality
2-D transthoracic echocardiography
Patients
500
Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/chemistry.differentiable-render-camouflage-data
DRC CARLA Multi-Vehicle Camouflage Dataset
This repository releases the audited synthetic data and geometry assets used to
study whether one differentiable vehicle-camouflage generator transfers across
vehicle shapes. The release preserves the original split manifests, collection
protocols, audit records, and SHA-256 checksums. It is intended for reproducible
adversarial-robustness research, including the analysis of negative results.
Release contents… See the full description on the dataset page: https://huggingface.co/datasets/bailuyucha/differentiable-render-camouflage-data.wavelet-lstm-camels-models
Wavelet-LSTM CAMELS Streamflow Models
A collection of 61,380 pre-trained LSTM models for daily streamflow forecasting across 620 USGS catchments from the CAMELS dataset.
Each catchment has 99 independently trained models:
33 wavelet filters × 3 lead times (1, 3, 5 days) = 99 wavelet-enhanced models
33 matching baseline models (same architecture, no wavelet transform)
Models are designed to be ensembled across wavelets for robust predictions with uncertainty estimates.… See the full description on the dataset page: https://huggingface.co/datasets/johnswyou/wavelet-lstm-camels-models.biology
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We provide… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/biology.seta-env-harborCambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.SETA-Env
SETA-Env
SETA-Env is an open-source verifiable RL terminal environment dataset for community training and evaluation.
This release contains two top-level subsets:
SETA_Synth: synthesized tasks
SETA_Evolve: evolved variants of terminal-agent tasks
The current release contains 4567 environments:
SETA_Synth: 3255
SETA_Evolve: 1312
What Is Included
Each task is packaged as a self-contained Harbor-style task directory with the files needed to run the task, build… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/SETA-Env.Cambrian-S-3M
Cambrian-S-3M
TLDR: This is a collection of open-source video instruction tuning data used in Cambrian-S's third training stage.
Overview
Cambrian-S-3M combines three video instruction datasets:
Cambrian-S-3M
LLaVA-Video-178K
LLaVA-Hound (ShareGPTVideo)
Prerequisites
Hugging Face CLI: pip install -U "huggingface_hub[cli]==0.36.0"
Sufficient disk space (~5 TB recommended)
hf command should be available after installing huggingface_hub
Setup… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-S-3M.math
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Math dataset is composed of 50K problem-solution pairs obtained using GPT-4. The dataset problem-solutions pairs generating from 25 math topics, 25 subtopics for each topic and 80 problems for each "topic,subtopic" pairs.
We provide the… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/math.Articraft-10KThis repository contains the 10k articulated 3D objects (in URDF format) from Articraft-10K.
Articraft-10K is a large-scale articulated 3D dataset generated by the Articraft agent.
arctic-camtraj-comparison-20260915
ARCTIC:六段视频与相机轨迹评测审计
在线同步查看 · 新版公开结果目录 · 离线 ZIP
本次新增五段,共六段视频;每格左侧视频 + MANO,右侧点云 + 相机与手腕轨迹。新增当前片段分数、GT/预测相机轨迹、逐采样点位置误差与仅相机引起的世界手腕差异。后者不是真实 world hand MPJPE。
原 36 段因手部 GT 覆盖条件排除了一些相机可评测片段,本次将剩余 27 段单独冻结并实际跑完七个条件。合计 63 段 / 30 条录制 / 单一被试 s05。441 条轨迹已用独立 evo 实现核验,提供录制级置信区间及多重比较检查。
可复算的单人短片段结果,不足以证明能全面替换现有 ViPE。 标定 ViPE / 30 Hz 得到 K + 预测深度及更多输入帧;不是同输入的纯模型比较,也不是完整实测深度生产管线。DA3 / VGGT 原生单位无米制保证,Sim3 校正尺度后误差不能称为原生米制精度。
完整审计报告 · 63 段逐片段指标
原始 36 段与单视频交付目录保留。现有 V9P3R1… See the full description on the dataset page: https://huggingface.co/datasets/yangzijing/arctic-camtraj-comparison-20260915.CameraBench
📷 CameraBench: Towards Understanding Camera Motions in Any Video
SfMs and VLMs performance on CameraBench: Generative VLMs (evaluated with VQAScore) trail classical SfM/SLAM in pure geometry, yet they outperform discriminative VLMs that rely on CLIPScore/ITMScore and—even better—capture scene‑aware semantic cues missed by SfM
After simple supervised fine‑tuning (SFT) on ≈1,400 extra annotated clips, our 7B Qwen2.5‑VL doubles its AP, outperforming the current best… See the full description on the dataset page: https://huggingface.co/datasets/syCen/CameraBench.koch_wrist_cam_depth_1tbench-tasks_migratedITLP-Campus-Outdoor🌳 ITLP Campus Outdoor is a multimodal dataset for Place Recognition research in diverse university campus environments. Captured by a mobile robot with front and back RGB cameras and a 3D LiDAR, it covers 3.6 km of outdoor paths across different seasons (winter and spring) and times of day (day, night, twilight). The dataset includes synchronized LiDAR point clouds, RGB images, semantic segmentation masks, and natural language scene descriptions for many of the frames. Semantic masks and text… See the full description on the dataset page: https://huggingface.co/datasets/OPR-Project/ITLP-Campus-Outdoor.BabyLMDataset for the shared baby language modeling task.
The goal is to train a language model from scratch on this data which represents
roughly the amount of text and speech data a young child observes.IDLE-OO-Camera-Traps
Dataset Card for IDLE-OO Camera Traps
IDLE-OO Camera Traps is a 5-dataset benchmark of camera trap images from the Labeled Information Library of Alexandria: Biology and Conservation (LILA BC) with a total of 2,586 images for species classification. Each of the 5 benchmarks is balanced to have the same number of images for each species within it (between 310 and 1120 images), representing between 16 and 39 species.
Supported Tasks and Leaderboards
Image… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/IDLE-OO-Camera-Traps.CAMELYON16
CAMELYON16
1. Tổng quan
CAMELYON16 là dataset ảnh mô bệnh học toàn tiêu bản (WSI) hạch bạch huyết canh gác (sentinel lymph node) của bệnh nhân ung thư vú, thu thập tại 2 trung tâm ở Hà Lan (Radboud University Medical Center và University Medical Center Utrecht). Bài toán chính là phân loại nhị phân cấp-slide: phát hiện có/không có di căn ung thư trong hạch (tumor/normal). Dataset gốc gồm 400 WSI (270 training, 130 testing).
Nguồn dữ liệu: AWS Open Data… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON16.camus
CAMUS - 2-D Echocardiographic Ultrasound Dataset
This dataset is a zea-format (HDF5) conversion of the
CAMUS
dataset for multi-structure segmentation in 2-D echocardiography.
Property
Value
Modality
2-D transthoracic echocardiography
Patients
500
Views
2-chamber (2CH) and 4-chamber (4CH) apical
Splits
train (1-400), val (401-450), test (451-500)
Conversion
This dataset was downloaded, converted to zea format, and uploaded using the
zea data… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus.cbt
Dataset Card for CBT
Dataset Summary
The Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context. The CBT is built from books that are freely available.
This dataset contains four different configurations:
V: where the answers to the questions are verbs.
P: where the answers to the questions are pronouns.
NE: where the answers to the questions are named entities.
CN: where the answers to the questions are… See the full description on the dataset page: https://huggingface.co/datasets/cam-cst/cbt.seta-env-evolcamera_pizza_additionalDDD-Cambodia-khmer-speech-dataset-v0-2-0CameraBenchProrobotwin2.0-egowam-flow-cam_high-subset
RoboTwin 2.0 cam_high EgoWAM-style 3D Flow — snapshot subset (8,522 episodes)
Query-based 3D motion-flow sidecar labels for the RoboTwin 2.0 LeRobot dataset
(yuanty/robotwin2.0-fastwam), generated with a pretrained 3D point tracker
(Track4World, DA3 backbone, metric-scale mode) from RGB only — EgoWAM-style
(arXiv 2607.08436 §4.2 conventions). This is a frozen snapshot of an in-progress
full-dataset run (27,500 episodes); see snapshot_manifest.json for the exact
episode list and… See the full description on the dataset page: https://huggingface.co/datasets/kimtaey/robotwin2.0-egowam-flow-cam_high-subset.spatial-cambrian-ssr-geothinker-release
