datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anomaly_detection_cmsl1t
Trigger Anomaly Detection for New Physics at the Large Hadron Collider
This dataset is a mirror of the Zenodo record: https://doi.org/10.5281/zenodo.21787779
This dataset contains Level-1 Trigger objects from the CMS experiment at the CERN Large
Hadron Collider, assembled for research on unsupervised anomaly detection in the
trigger. The goal of unsupervised anomaly detection in this context is the discovery of
new physics. This data set does not contain new physics. It is meant… See the full description on the dataset page: https://huggingface.co/datasets/CERN/anomaly_detection_cmsl1t.spider
Dataset Card for "spider"
More Information needed
vqa-cmsv-benchmark
VQA-CMSV Benchmark Data Package
This repository contains annotation splits for VQA v2-CMSV, GQA-CMSV, and VG-CMSV, plus patch-mask NPZ files used for mask supervision experiments.
Contents
data/vqa_v2_cmsv/train.json, data/vqa_v2_cmsv/val.json, data/vqa_v2_cmsv/test.json
data/gqa_cmsv/train.jsonl, data/gqa_cmsv/val.jsonl, data/gqa_cmsv/test.jsonl
data/vg_cmsv/train.jsonl, data/vg_cmsv/val.jsonl, data/vg_cmsv/test.jsonl
masks/vqa_v2_cmsv_masks.npz
masks/gqa_cmsv_masks.npz… See the full description on the dataset page: https://huggingface.co/datasets/as-benchmark-artifacts/vqa-cmsv-benchmark.cms-desynpuf-insurance-claims
CMS DE-SynPUF Insurance Claims (Rendered EOB Documents)
Pipeline-ready insurance_claim evaluation corpus derived from the CMS
2008-2010 Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF), Sample 1.
One row = one Medicare claim rendered as a plain-text EOB-style document with
ground-truth extraction fields aligned to the llm-mailroom
InsuranceClaimExtraction schema.
Provenance
Source: CMS DE-SynPUF Sample 1 (fully synthetic; "very limited inferential… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/cms-desynpuf-insurance-claims.kl3m-data-dotgov-www.cms.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cms.gov.regex-pattern-generation-with-verified-match-sets-cmskdvm7
Regex Pattern Generation with Verified Match Sets
A dataset of regex pattern generation with verified match sets examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Regex
Framework: Community
License: CC-BY-4.0… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/regex-pattern-generation-with-verified-match-sets-cmskdvm7.personal-blog-cmspython-execution-trace-output-prediction-cmskdimp
Python Execution Trace & Output Prediction
Self-contained Python programs paired with a concrete function call and the exact runtime output. Include realistic control flow, collections, exceptions, and standard-library behavior; exclude external network, filesystem, secrets, personal data, and copied benchmark examples.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-execution-trace-output-prediction-cmskdimp.cms-nursing-home-quality-and-ownership
Canonical source: https://www.fastdol.com/datasets/cms-nursing-home-quality-and-ownership
US Nursing Home Quality, Ownership & Enforcement (CMS)
14,693 Medicare/Medicaid-certified nursing homes · one row per facility · 66 columns · CMS Care Compare, refreshed 2026-08-19 · certifications 1967-01-01 → surveys through 2026-06-26.
Every certified U.S. nursing home in a single table: CMS 5-star quality ratings, payroll-based staffing and turnover, corporate ownership and chain… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/cms-nursing-home-quality-and-ownership.build-competitive-programming-problem-dataset-cmsokff2
Build Competitive Programming Problem Dataset
Each item is a build competitive programming problem dataset example providing Problem statement, Constraints, Reference solution, Test cases, Time limit (ms). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: Python… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/build-competitive-programming-problem-dataset-cmsokff2.graphql-schema-resolver-execution-validation-cmskdigf
GraphQL Schema Resolver Execution Validation
Each item is a graphql schema resolver execution validation example providing Schema (SDL), Resolver map, Sample query, Expected response (JSON). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: JavaScript
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/graphql-schema-resolver-execution-validation-cmskdigf.python-data-transformation-scripts-cmskdijz
Python Data Transformation Scripts
Each item is a python data transformation scripts example providing Input sample, What the transform should do, Transformation script, Expected output. Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-data-transformation-scripts-cmskdijz.extract-metrics-from-log-fixtures-cmskdlip
Extract Metrics from Log Fixtures
A dataset of extract metrics from log fixtures examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework: Community
License: CC-BY-4.0
Contributors… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/extract-metrics-from-log-fixtures-cmskdlip.decompilebench-with-addressescms_iom_500
Dataset Summary
This dataset contains CMS information with local and national coverage document data sets (LCD & NCD),as Coverage Articles and [Internet-Only Manuals (IOMs)(https://www.cms.gov/medicare/regulations-guidance/manuals/internet-only-manuals-ioms)
A list of Current LCDS, NCDs and Articles is obrained from Medicare Coverage Database.
The data itself was obtainted by scrapping the urls and extracting data from the pdf files listed in current articles and current lcds… See the full description on the dataset page: https://huggingface.co/datasets/evekhm/cms_iom_500.TiMoS_alloys_CMS2021
Cite this dataset Silva, A., Polcar, T., and Kramer, D. TiMoS alloys CMS2021. ColabFit, 2023. https://doi.org/10.60732/864a2df0
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jn819esw58ah_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/TiMoS_alloys_CMS2021.coherent-forms-1040-cms1500-i9
SymageDocs — Coherent US Tax / Health / Employment Forms (FUNSD)
A fully synthetic document-AI training set: three US forms — IRS Form 1040,
CMS-1500, and USCIS Form I-9 — filled from the same synthetic
identity, so name / SSN / address / employer flow consistently across all
three renderings. Each page ships with FUNSD ground truth (word boxes,
entity labels, key–value linking) plus a LayoutLMv3-ready token/bbox/tag view.
3,000 page-level image + annotation rows (train 2,400 /… See the full description on the dataset page: https://huggingface.co/datasets/Symage/coherent-forms-1040-cms1500-i9.LLM_artist_lyricsTiO2_CMS2016
Cite this dataset Artrith, N., and Urban, A. TiO2 CMS2016. ColabFit, 2023. https://doi.org/10.60732/861c6a25
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_kvjft3au55qb_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org
Dataset Name… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/TiO2_CMS2016.Zn_MTP_CMS2023
Cite this dataset Mei, H., Cheng, L., Chen, L., Wang, F., Li, J., and Kong, L. Zn MTP CMS2023. ColabFit, 2024. https://doi.org/10.60732/54902e18
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_58y020ce6b6j_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Zn_MTP_CMS2023.cms_iom_3000
Dataset Summary
This dataset contains CMS information with local and national coverage document data sets (LCD & NCD),as Coverage Articles and [Internet-Only Manuals (IOMs)(https://www.cms.gov/medicare/regulations-guidance/manuals/internet-only-manuals-ioms)
A list of Current LCDS, NCDs and Articles is obrained from Medicare Coverage Database.
The data itself was obtainted by scrapping the urls and extracting data from the pdf files listed in current articles and current lcds… See the full description on the dataset page: https://huggingface.co/datasets/evekhm/cms_iom_3000.cms-glp1-data
The GLP-1 Spending Explosion (2017-2025)
Everyone's heard about Ozempic, Wegovy, and Mounjaro by now. But how much is Medicare actually spending on these drugs, and how fast is it growing? That's what this dataset is about.
I pulled together state-by-state Medicare Part D spending data on GLP-1 medications going back to 2017, when Ozempic first got FDA approval. The dataset covers all 50 states and territories through 2023 at the state level, plus provisional 2024 and 2025 national… See the full description on the dataset page: https://huggingface.co/datasets/sbreitenbach/cms-glp1-data.JiRack-Pretrain-DatasetThis dataset was created with a strong focus on corporate AI. Its main goal is to help banks, fintech companies, and enterprises protect sensitive data and stay compliant with strict U.S. privacy laws that prohibit the use of public online and cloud-based AI services.
Multi-Domain High-Quality Corpus
A carefully curated high-quality multilingual dataset used to pretrain the JiRack model with modern data.
So JiRack Tokenizers stand out because they’ve been trained on premium… See the full description on the dataset page: https://huggingface.co/datasets/CMSManhattan/JiRack-Pretrain-Dataset.CoNbV_CMS2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. CoNbV CMS2019. ColabFit, 2023. https://doi.org/10.60732/f2c623f1
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_sn623uhg2d1b_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CoNbV_CMS2019.CMSBCuPd_CMS2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. CuPd CMS2019. ColabFit, 2023. https://doi.org/10.60732/1058e01c
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_0ry6z1j8mi8c_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CuPd_CMS2019.AlNiTi_CMS_2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. AlNiTi CMS 2019. ColabFit, 2023. https://doi.org/10.60732/7b56ca82
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_dtjyh96dypuu_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/AlNiTi_CMS_2019.cms_bert-vocab
Dataset Card for "cms_bert-vocab"
More Information needed
medicaid-cms-64-new-adult-group-expenditures
Medicaid CMS-64 New Adult Group Expenditures
Description
This dataset reports summary level expenditure data associated with the new adult group established under the Affordable Care Act. These state expenditures are reported through the federal Medicaid Budget and Expenditure System (MBES).
Notes:
“VIII GROUP” is also known as the “New Adult Group.”
The VIII Group is only applicable for states that have expanded their Medicaid programs by adopting the VIII Group. VIII… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/medicaid-cms-64-new-adult-group-expenditures.ocr-bert_cms-vocab
Dataset Card for "ocr-bert_cms-vocab"
More Information needed
