datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conflux-chest-ct
CONFLUX Chest-CT
200,000 synthetic 3D chest CT volumes with structured abnormality and demographic labels, generated by CONFLUX.
Released with the paper CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training.
Paper (arXiv) •
Model •
Code — coming soon
About
CONFLUX is a conditional 3D latent generative model for chest CT: a VAE tokenizer
compresses each volume into a compact 16-channel latent, a… See the full description on the dataset page: https://huggingface.co/datasets/gevaertlab/conflux-chest-ct.forecast-snapshots-metaculus-6f1cdfd9b3
Forecast Snapshot Dataset
Source: metaculus
Data Hash: 6f1cdfd9b3
Dataset Description
This dataset contains snapshots of prediction market data for AI agent evaluation.
Format
Each row represents a market snapshot at a specific point in time, including:
Current community prediction
Predictions 1 day, 3 days, 1 week, 2 weeks, 1 month later (or resolution)
Final resolution (ground truth)
Market metadata (question, close time, etc.)
Question Types… See the full description on the dataset page: https://huggingface.co/datasets/chestnutforty/forecast-snapshots-metaculus-6f1cdfd9b3.robomind_benchmark1_0_release_franka_3rgb_pick_apple_into_chest
benchmark1_0_release_franka_3rgb_pick_apple_into_chest
This dataset converts the Robomain format uniformly into LeRobot V3.0.
Dataset Statistics
本体: franka_3rgb
末端执行器: 夹爪
任务平台显示版: 将苹果捡起放入箱子
total_episodes: 2
total_tasks: 1
size: 17.0M
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── images
│ └── observation.images.camera_top
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/robomind_benchmark1_0_release_franka_3rgb_pick_apple_into_chest.huac_labeled_chest_xray_reports_c5_v1forecast-snapshots-metaculus-2cc65706d0
Forecast Snapshot Dataset
Source: metaculus
Data Hash: 2cc65706d0
Dataset Description
This dataset contains snapshots of prediction market data for AI agent evaluation.
Format
Each row represents a market snapshot at a specific point in time, including:
Current community prediction
Predictions 1 day, 3 days, 1 week, 2 weeks, 1 month later (or resolution)
Final resolution (ground truth)
Market metadata (question, close time, etc.)
Question Types… See the full description on the dataset page: https://huggingface.co/datasets/chestnutforty/forecast-snapshots-metaculus-2cc65706d0.forecast-snapshots-kalshi_events-768472771c
Forecast Snapshot Dataset
Source: kalshi_events
Data Hash: 768472771c
Dataset Description
This dataset contains snapshots of prediction market data for AI agent evaluation.
Format
Each row represents a market snapshot at a specific point in time, including:
Current community prediction
Predictions 1 day, 3 days, 1 week, 2 weeks, 1 month later (or resolution)
Final resolution (ground truth)
Market metadata (question, close time, etc.)
Question Types… See the full description on the dataset page: https://huggingface.co/datasets/chestnutforty/forecast-snapshots-kalshi_events-768472771c.chest2vec_labels
CT-RATE Findings — Chest Imaging Leaf Labels
Chest-CT findings from the CT-RATE
dataset. Each row maps an original findings report → a section-structured refined
version, plus a 137-label ternary multi-label annotation over a chest-imaging taxonomy.
Rows: 23,614 unique CT-RATE findings reports (one row per report)
Splits (report-text-level de-duplicated): train 20,648 / valid 1,483 / test 1,483
Labels: 137 leaf labels (106 clinical + 31 other). Full taxonomy, definitions and… See the full description on the dataset page: https://huggingface.co/datasets/chest2vec/chest2vec_labels.gpt-oss-20b-red-teaming-evals
GPT-OSS-20b Red-Teaming Mass Evaluation
Overview
This dataset contains 16,181 prompt-response evaluation pairs from OpenAI's gpt-oss-20b model, generated as part of a large-scale red-teaming effort for the Kaggle Red-Teaming Challenge.
The evaluations are sourced from 15+ distinct public red-teaming and safety datasets. Each record includes the original prompt, the model's response, token counts, a harm category classification, the source dataset, and binary flags for… See the full description on the dataset page: https://huggingface.co/datasets/ChestnutKurisu/gpt-oss-20b-red-teaming-evals.NIH_Chest_XRay_Local_Balancedaadigupta1601_chest-x-ray-pneumonia-numerical-feature-dataset
Chest X-Ray Pneumonia Numerical Feature Dataset
Validated Numerical Features Extracted from Chest X-Ray Image
Dataset Info
Source: Kaggle
Original Size: 1.10 MB
Kaggle Downloads: 418
Files: 3
Files
test_features.csv
train_features.csv
val_features.csv
Mirrored from Kaggle
huac_labeled_chest_xray_reports_no_synthetic_v1huac_labeled_chest_xray_reports_v6Filtered_NIH-Chest-Xraychester_bennington_30tracks_dataset137k_chest_xray_label
