datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anomaly_detection_cmsl1t
Trigger Anomaly Detection for New Physics at the Large Hadron Collider
This dataset is a mirror of the Zenodo record: https://doi.org/10.5281/zenodo.21787779
This dataset contains Level-1 Trigger objects from the CMS experiment at the CERN Large
Hadron Collider, assembled for research on unsupervised anomaly detection in the
trigger. The goal of unsupervised anomaly detection in this context is the discovery of
new physics. This data set does not contain new physics. It is meant… See the full description on the dataset page: https://huggingface.co/datasets/CERN/anomaly_detection_cmsl1t.global-medicines-atlas-cms-partd-20260827
CMS Part D public source corpus
Exact official public resources for the CMS Quarterly Prescription Drug Plan
Formulary, Pharmacy Network, and Pricing Information catalogue through Q2 2026,
plus the three catalogue resources for Medicare Part D Spending by Drug for the
2024 reporting period.
CMS is the source. The formulary resources remain subject to the CMS Agreement
for Use. Uses must be accurate and non-misleading, preserve source dates,
specification changes, suppression… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/global-medicines-atlas-cms-partd-20260827.spider
Dataset Card for "spider"
More Information needed
dataset_20250901_A_60eps
dataset_20250901_A_60eps
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
vidaio-cmscms-desynpuf-insurance-claims
CMS DE-SynPUF Insurance Claims (Rendered EOB Documents)
Pipeline-ready insurance_claim evaluation corpus derived from the CMS
2008-2010 Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF), Sample 1.
One row = one Medicare claim rendered as a plain-text EOB-style document with
ground-truth extraction fields aligned to the llm-mailroom
InsuranceClaimExtraction schema.
Provenance
Source: CMS DE-SynPUF Sample 1 (fully synthetic; "very limited inferential… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/cms-desynpuf-insurance-claims.cms-medicare-database
CMS Medicare Physician & Other Supplier Database
A clean, queryable DuckDB database built from the CMS Medicare Physician & Other Practitioners Public Use Files -- provider-level Medicare Part B claims data from CY2012 through CY2023.
132,947,347 rows across 3 tables covering what every physician billed, what Medicare paid, and how many services and beneficiaries per NPI per HCPCS code.
Built with cms-medicare-database.
Quick Start
DuckDB CLI… See the full description on the dataset page: https://huggingface.co/datasets/Nason/cms-medicare-database.place-food-in-bowlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 72000,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cmsng2001/place-food-in-bowl.vqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.regex-pattern-generation-with-verified-match-sets-cmskdvm7
Regex Pattern Generation with Verified Match Sets
A dataset of regex pattern generation with verified match sets examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Regex
Framework: Community
License: CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/regex-pattern-generation-with-verified-match-sets-cmskdvm7.personal-blog-cmspython-execution-trace-output-prediction-cmskdimp
Python Execution Trace & Output Prediction
Self-contained Python programs paired with a concrete function call and the exact runtime output. Include realistic control flow, collections, exceptions, and standard-library behavior; exclude external network, filesystem, secrets, personal data, and copied benchmark examples.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-execution-trace-output-prediction-cmskdimp.PAXRay
PAXRay Dataset
PAXRay (Projected dataset for the segmentation of Anatomical structures in X-Ray data) is a synthetic dataset designed to facilitate research on anatomical segmentation in X-Ray imagery. The dataset contains projections of the RibFrac CT dataset onto a 2D plane to simulate realistic X-Ray data. It includes multi-label segmentation masks for fine-grained and hierarchical anatomical labeling.
For an extended version of this dataset, check out PAXRay++.
For models for… See the full description on the dataset page: https://huggingface.co/datasets/cmseibold/PAXRay.classifier_setscms-nursing-home-quality-and-ownership
Canonical source: https://www.fastdol.com/datasets/cms-nursing-home-quality-and-ownership
US Nursing Home Quality, Ownership & Enforcement (CMS)
14,693 Medicare/Medicaid-certified nursing homes · one row per facility · 66 columns · CMS Care Compare, refreshed 2026-08-19 · certifications 1967-01-01 → surveys through 2026-06-26.
Every certified U.S. nursing home in a single table: CMS 5-star quality ratings, payroll-based staffing and turnover, corporate ownership and chain… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/cms-nursing-home-quality-and-ownership.EveNet-GridStudy-CMSOpenData
📄 Citation
If you use this dataset, please cite:
@article{zhang2026evenet,
title={EveNet: A Foundation Model for Particle Collision Data Analysis},
author={Zhang, Yulei and others},
journal={arXiv preprint arXiv:2601.17126},
year={2026}
}
kl3m-data-dotgov-www.cms.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cms.gov.build-competitive-programming-problem-dataset-cmsokff2
Build Competitive Programming Problem Dataset
Each item is a build competitive programming problem dataset example providing Problem statement, Constraints, Reference solution, Test cases, Time limit (ms). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Python… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/build-competitive-programming-problem-dataset-cmsokff2.PAX-RayPlusPlus
PAX-Ray++ Dataset
The PAX-Ray++ Dataset is a high-quality dataset designed to facilitate segmentation tasks for anatomical structures in chest radiographs. By leveraging pseudo-labeled thorax CT scans projected onto a 2D plane, this dataset provides fine-grained annotations resembling traditional X-ray imaging. This enables the development and evaluation of models tailored to anatomical segmentation in medical imaging.
Key Features
Large Dataset: Contains 7,377… See the full description on the dataset page: https://huggingface.co/datasets/cmseibold/PAX-RayPlusPlus.graphql-schema-resolver-execution-validation-cmskdigf
GraphQL Schema Resolver Execution Validation
Each item is a graphql schema resolver execution validation example providing Schema (SDL), Resolver map, Sample query, Expected response (JSON). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: JavaScript
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/graphql-schema-resolver-execution-validation-cmskdigf.c_ms_girlsfrontline
Dataset of c_ms/C-MS/C-MS (Girls' Frontline)
This is the dataset of c_ms/C-MS/C-MS (Girls' Frontline), containing 117 images and their tags.
The core tags of this character are black_hair, long_hair, red_eyes, bangs, mole, mole_under_eye, very_long_hair, breasts, messy_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/c_ms_girlsfrontline.python-data-transformation-scripts-cmskdijz
Python Data Transformation Scripts
Each item is a python data transformation scripts example providing Input sample, What the transform should do, Transformation script, Expected output. Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-data-transformation-scripts-cmskdijz.decompilebench-with-addressesextract-metrics-from-log-fixtures-cmskdlip
Extract Metrics from Log Fixtures
A dataset of extract metrics from log fixtures examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework: Community
License: CC-BY-4.0
Contributors… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/extract-metrics-from-log-fixtures-cmskdlip.cmsaf-ocean-water-flux-1987-2014cms-hospitals
CMS Hospital General Information
Every Medicare-certified US hospital (~5,432) with CMS overall star rating (1-5), ownership type, emergency-services flag, and quality-measure counts across mortality, safety, readmissions, patient experience.
Live API
This dataset is served via a live REST API at api.ai-analytics.org. The card you're reading exists so HuggingFace's index can route AI agents + researchers to the canonical source.
API endpoint:… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/cms-hospitals.cms_federal_medicare
Dataset Card for US Dialysis Facilities
Dataset Details
The dataset includes a wide range of metrics, such as Five Star ratings, addresses, city/town, state, and various statistical measures related to the quality and outcomes of the facilities.
Dataset Description
The "DFC_FACILITY.csv" dataset contains information about dialysis facilities, including certification, ratings, locations, and various performance measures.
Curated by: Centers for Medicare and… See the full description on the dataset page: https://huggingface.co/datasets/stigsfoot/cms_federal_medicare.LLM_artist_lyricsTiMoS_alloys_CMS2021
Cite this dataset Silva, A., Polcar, T., and Kramer, D. TiMoS alloys CMS2021. ColabFit, 2023. https://doi.org/10.60732/864a2df0
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jn819esw58ah_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/TiMoS_alloys_CMS2021.Marker_pickup_piper
Marker_pickup_piper
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
