datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anomaly_detection_cmsl1t
Trigger Anomaly Detection for New Physics at the Large Hadron Collider
This dataset is a mirror of the Zenodo record: https://doi.org/10.5281/zenodo.21787779
This dataset contains Level-1 Trigger objects from the CMS experiment at the CERN Large
Hadron Collider, assembled for research on unsupervised anomaly detection in the
trigger. The goal of unsupervised anomaly detection in this context is the discovery of
new physics. This data set does not contain new physics. It is meant… See the full description on the dataset page: https://huggingface.co/datasets/CERN/anomaly_detection_cmsl1t.spider
Dataset Card for "spider"
More Information needed
place-food-in-bowlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 72000,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cmsng2001/place-food-in-bowl.kl3m-data-dotgov-www.cms.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.cms.gov.decompilebench-with-addressesTiMoS_alloys_CMS2021
Cite this dataset Silva, A., Polcar, T., and Kramer, D. TiMoS alloys CMS2021. ColabFit, 2023. https://doi.org/10.60732/864a2df0
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jn819esw58ah_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/TiMoS_alloys_CMS2021.coherent-forms-1040-cms1500-i9
SymageDocs — Coherent US Tax / Health / Employment Forms (FUNSD)
A fully synthetic document-AI training set: three US forms — IRS Form 1040,
CMS-1500, and USCIS Form I-9 — filled from the same synthetic
identity, so name / SSN / address / employer flow consistently across all
three renderings. Each page ships with FUNSD ground truth (word boxes,
entity labels, key–value linking) plus a LayoutLMv3-ready token/bbox/tag view.
3,000 page-level image + annotation rows (train 2,400 /… See the full description on the dataset page: https://huggingface.co/datasets/Symage/coherent-forms-1040-cms1500-i9.TiO2_CMS2016
Cite this dataset Artrith, N., and Urban, A. TiO2 CMS2016. ColabFit, 2023. https://doi.org/10.60732/861c6a25
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_kvjft3au55qb_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.
https://materials.colabfit.org
Dataset Name… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/TiO2_CMS2016.Zn_MTP_CMS2023
Cite this dataset Mei, H., Cheng, L., Chen, L., Wang, F., Li, J., and Kong, L. Zn MTP CMS2023. ColabFit, 2024. https://doi.org/10.60732/54902e18
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_58y020ce6b6j_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Zn_MTP_CMS2023.cms_rules1CoNbV_CMS2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. CoNbV CMS2019. ColabFit, 2023. https://doi.org/10.60732/f2c623f1
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_sn623uhg2d1b_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CoNbV_CMS2019.CMSBCuPd_CMS2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. CuPd CMS2019. ColabFit, 2023. https://doi.org/10.60732/1058e01c
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_0ry6z1j8mi8c_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/CuPd_CMS2019.AlNiTi_CMS_2019
Cite this dataset Gubaev, K., Podryabinkin, E. V., Hart, G. L., and Shapeev, A. V. AlNiTi CMS 2019. ColabFit, 2023. https://doi.org/10.60732/7b56ca82
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_dtjyh96dypuu_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/AlNiTi_CMS_2019.cms_bert-vocab
Dataset Card for "cms_bert-vocab"
More Information needed
ocr-bert_cms-vocab
Dataset Card for "ocr-bert_cms-vocab"
More Information needed
kl3m-filter-data-dotgov-www.cms.govcms_11cms-icd10-categorical
Dataset Card for "cms-icd10-categorical"
More Information needed
cms-nppes
CMS NPPES — historical annual snapshots
Historical CMS National Plan and Provider Enumeration System (NPPES) files ingested from the NBER NPPES mirror (https://data.nber.org/npi) and staged as annual, analysis-ready Parquet (zstd-compressed). The pipeline crawls the NBER directory, downloads the source files, converts and unifies them, then validates the National Provider Identifier (NPI) via Luhn check, uniqueness, and a cross-check against the raw CMS extract before promoting… See the full description on the dataset page: https://huggingface.co/datasets/mauricedalton/cms-nppes.CMSC848K-House
